Skip to content

Repository files navigation

Detection-Rule Leakage and Trust Camouflage

Zenodo DOI

Research package

Field Record
Author Michael Zot
ORCID 0009-0001-9194-938X
Official OSF DOI https://doi.org/10.17605/OSF.IO/XG82Q
PhilPapers / PhilArchive https://philpapers.org/rec/ZOTDLA
NCBI My Bibliography https://www.ncbi.nlm.nih.gov/myncbi/1Vch2ALud0Sk3I/bibliography/public/
Zenodo GitHub Archive DOI https://doi.org/10.5281/zenodo.20779751
GitHub Materials Mirror https://github.com/mikecreation/detection-rule-leakage-trust-camouflage

This repository mirrors the research materials for Detection-Rule Leakage and Trust Camouflage, a function-over-form framework for studying adversarial adaptation in open-audience mental-health communication.

The project examines a specific failure mode in public abuse-detection education: the same detection rules that help victims, bystanders, clinicians, institutions, and AI systems identify harmful behavior may also teach motivated actors what signals to stop displaying. The framework argues that detection should focus less on trusted language and more on behavioral function over time: repair, accountability, evidence handling, consequences, and repeated pattern.

Official archive

The official peer-review package and study materials are archived on OSF:

https://doi.org/10.17605/OSF.IO/XG82Q

The OSF archive should be treated as the canonical citation source. This GitHub repository is a public mirror for easier access, repository inspection, code review, and collaboration.

PhilPapers / PhilArchive record

This work is listed on PhilPapers / PhilArchive:

https://philpapers.org/rec/ZOTDLA

This record provides an additional philosophy and cross-disciplinary visibility layer for the manuscript.

NCBI bibliography record

This work is listed in Michael Zot’s public NCBI My Bibliography:

https://www.ncbi.nlm.nih.gov/myncbi/1Vch2ALud0Sk3I/bibliography/public/

This record provides an additional public bibliography layer through the National Library of Medicine account system. It does not mean the work has been indexed in PubMed.

Archived GitHub release

This GitHub materials mirror is archived on Zenodo:

https://doi.org/10.5281/zenodo.20779751

The OSF archive remains the canonical citation source. The Zenodo DOI preserves the GitHub release for repository inspection, long-term access, and indexing.

Repository contents

This repository includes:

  • Main scholarly manuscript
  • Trust Camouflage Function Score coding protocol
  • H2 preregistration materials
  • H2 vignette stimuli
  • Qualtrics / jsPsych pilot materials
  • TC-LaunderBench materials for LLM rhetorical laundering tests
  • TCFS coder sheet and coder instructions
  • Inter-rater reliability analysis script
  • Empirical validation roadmap
  • MIPU reversal cost protocol
  • H3 composure credibility protocol
  • Literature-gap search materials
  • Manifest files
  • Optional complete archive: All Research Materials.zip

Folder map

Folder / file Purpose
Scholarly Manuscript/ Main manuscript and related paper files
TCFS/ Trust Camouflage Function Score coding materials
H2 Pilot/ H2 preregistration, vignette, Qualtrics, and jsPsych materials
H3 Pilot/ Composure credibility protocol materials
TC LaunderBench/ LLM rhetorical laundering benchmark materials
Validation/ Validation roadmap and supporting materials
Literature Gap/ Literature-gap search worksheet and related materials
Admin/ Manifest and administrative files
All Research Materials.zip Optional one-click archive of the full package
CITATION.cff GitHub citation metadata
README.md Repository overview

Core constructs

Detection-Rule Leakage

Detection-Rule Leakage describes the risk that public detection advice can teach two audiences at once: victims learn what to notice, while some harmful actors learn what signs to stop showing.

Trust Camouflage

Trust Camouflage describes the use of socially trusted language, including therapy-speak, boundary language, trauma language, safety language, or accountability language, to appear safe while preserving harmful behavioral function.

Function-over-form detection

The framework argues that language alone is insufficient. The central question is what the communication does over time:

  • Does it permit repair?
  • Does it preserve accountability?
  • Does it keep evidence available?
  • Does it allow consequences?
  • Does the pattern repeat?

Locked H2 design

H2 uses a 4-condition design:

  1. Blunt
  2. Neutral
  3. Therapeutic Mild
  4. Therapeutic Strong

The confirmatory primary contrast is:

Therapeutic Mild < Blunt on Accountability Composite.

Therapeutic Strong is exploratory only and is used for dose-response and ceiling-effect calibration.

Research-use guardrail

This is a pre-validation research package. The Trust Camouflage Function Score coding protocol is not a clinical instrument, diagnostic tool, custody tool, HR tool, or relationship verdict.

The materials are intended for research use, peer review, and controlled empirical testing.

Ethical-use notice

Do not use these materials to train manipulation models, generate interpersonal evasion scripts, create harassment workflows, or build tools that help people avoid accountability.

The intended use is measurement, peer review, controlled research, AI-safety testing, institutional audit design, and empirical validation.

License

Unless otherwise stated, research materials in this repository are licensed under:

Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International

The materials may be shared and adapted with attribution for noncommercial purposes, provided derivative works use the same license.

Code and analysis scripts may be reused for research and educational purposes with attribution.

Citation

Zot, M. (2026). Detection-Rule Leakage and Trust Camouflage: Peer-Review Package, Study Materials, and Validation Roadmap. OSF. https://doi.org/10.17605/OSF.IO/XG82Q

Related links

Releases

Packages

Contributors

Languages