Skip to content

Latest commit

 

History

History
759 lines (661 loc) · 45.5 KB

File metadata and controls

759 lines (661 loc) · 45.5 KB

Config, mass model, constants, schema, manifest

Part of the MuMDIA developer documentation (see docs/README.md).

Purpose

This document covers the mumdia-core crate: the shared vocabulary that every stage depends on but that contains no stage logic itself. Four concerns live here.

  1. Typed configuration (config.rs): one serde structure with per-stage sections and one strategy enum per algorithmic choice point. Config parses from JSON, rejects unknown keys, and validates on load so a misconfiguration fails at startup rather than producing a bogus run.
  2. Mass model (mass.rs + constants.rs): the single source of residue masses, physical constants, m/z conversion, ppm predicates, ProForma-lite peptidoform parsing, and b/y fragment generation. No other crate defines a mass constant or a ppm predicate.
  3. Frozen artifact schema ids (schema.rs): (logical name, version) pairs recorded in each artifact's report.json and in the manifest.json ArtifactRecord. They are provenance only: no stage reads them back, and Table::read (table.rs:207-225) does no version check, so a schema id is never used to detect a mismatch or invalidate a downstream artifact.
  4. Run manifest (manifest.rs): per-artifact provenance (content hash, producing stage, config hash) written once by the run orchestrator.

The crate is declared in lib.rs:6-13 (config, constants, error, manifest, mass, rejection, schema, types) and exposes version() from CARGO_PKG_VERSION (lib.rs:15-17).

Files

Path Role
rust/mumdia/crates/mumdia-core/src/config.rs All config structs + every strategy enum; Config::from_json, validate, apply_profile, canonical_json (1188 lines)
rust/mumdia/crates/mumdia-core/src/constants.rs Physical constants (PROTON, WATER, AMMONIA, ISOTOPE_SPACING), residue_mass, mass_to_mz, ppm predicates
rust/mumdia/crates/mumdia-core/src/mass.rs unimod_mass, IonType, Fragment, ParsedPeptidoform, parse_peptidoform, b/y fragment generation
rust/mumdia/crates/mumdia-core/src/schema.rs Frozen (name, version) ids for every artifact; PSMS_SCORED is v3, PSMS_COMPETED/PEPTIDE_QUANT/PROTEIN_GROUP_QUANT are v2, all others v1
rust/mumdia/crates/mumdia-core/src/manifest.rs Manifest + ArtifactRecord (provenance)
rust/mumdia/crates/mumdia-core/src/types.rs Peak, IsolationWindow, Label, Ms2Scan
rust/mumdia/crates/mumdia-core/src/rejection.rs RejectionReason ladder for the candidate-audit table
rust/mumdia/crates/mumdia-core/src/error.rs MassError, ConfigError (thiserror)
rust/mumdia/crates/mumdia-core/src/lib.rs Module wiring + version()
rust/mumdia/crates/mumdia-io/src/lib.rs record_artifact (builds ArtifactRecord by hashing the file), inspect
rust/mumdia/crates/mumdia-io/src/hash.rs blake3_file, blake3_str (content + config hashing)
rust/mumdia/crates/mumdia-io/src/report.rs ArtifactReport written as <artifact>.report.json

Inputs and outputs

mumdia-core produces no artifacts of its own. It defines the types the other crates read and write. Two things it owns appear on disk:

manifest.json (one per run, written at run.rs:500-501). Serialized Manifest; fields (manifest.rs:22-31):

Field Type Meaning
mumdia_version String CARGO_PKG_VERSION at build time (manifest.rs:36)
config_json String Fully-resolved config, from Config::canonical_json()
config_hash String blake3_str(canonical_json) (run.rs:89)
model_identities BTreeMap<String,String> rt_predictor, fragment_predictor, rescorer, feature_schema_id (run.rs:489-498)
artifacts BTreeMap<String,ArtifactRecord> one entry per produced artifact, keyed by logical name

ArtifactRecord (manifest.rs:9-20), one per artifact: logical_name, path, format (always "parquet"), schema_name, schema_version, rows, content_hash (blake3 of the file bytes), producing_stage, config_hash. Built by record_artifact (mumdia-io/src/lib.rs:20-39), which hashes the written Parquet file with blake3_file (hash.rs:8-20, streamed in 64 KiB chunks).

ArtifactReport / <artifact>.report.json (report.rs:11-24) is written alongside each artifact by its producing stage, not by core: logical_name, schema_name, schema_version, stage, rows, content_hash, params (resolved parameters), stats (key metrics), model_identity, elapsed_ms. write_for appends .report.json to the artifact path (report.rs:26-31).

Artifact schema ids (schema.rs:7-24)

Every id is a (&str, u32) constant. A stage stamps .0 and .1 into its ArtifactReport and ArtifactRecord. Both are provenance records only; nothing reads them back to validate a downstream read (Table::read, table.rs:207-225, does no version check, and write_table, table.rs:151-196, does not stamp the id into the Parquet file itself).

Constant Logical name Version
SPECTRA_MS1 spectra_ms1 1
SPECTRA_MS2 spectra_ms2 1
ISOLATION_WINDOWS isolation_windows 1
MS2_TO_MS1 ms2_to_ms1 1
PEPTIDES peptides 1
PEPTIDOFORMS peptidoforms 1
FRAGMENT_LIBRARY_PRECURSORS fragment_library_precursors 1
FRAGMENT_LIBRARY_FRAGMENTS fragment_library_fragments 1
SEED_PSMS seed_psms 1
RUN_WINDOWS run_windows 1
PSMS_EXTRACTED psms_extracted 1
CHROMATOGRAMS chromatograms 1
FEATURES features 1
PSMS_COMPETED psms_competed 2
PSMS_SCORED psms_scored 3
PEPTIDE_QUANT peptide_quant 2
PROTEIN_GROUP_QUANT protein_group_quant 2
FRAGMENT_QUANT fragment_quant 1

PSMS_SCORED v3 carries the identification apex/bounds through rescoring so quant can integrate the selected peak. Its column layout is written by the rescore stage and is the schema a downstream reader must expect:

Column Type Meaning
candidate_id U32 library candidate id
peptidoform Str ProForma-lite string
charge I32 precursor charge
label Str target / decoy
protein Str protein accession
base_peptide_id U32 interned stripped-peptide id (peptide-level q grouping)
apex_rt F64 identification apex used by downstream quant
elution_lo / elution_hi F64 identified elution bounds carried through
score F64 rescorer score
q_value F64 per-PSM q (pooled)
peptide_q_value F64 peptide-level q (global per base peptide)
protein_group Str protein-group key
pg_q_value F64 protein-group q
global_q_value F64 experiment-wide PSM q (== q_value for a single-run rescore)
prelim_score F64 pre-rescore feature-stage score
source U32 run index (0 for single-run; index into --competed for experiment-wide)
run_psm_q F64 per-run PSM FDR
experiment_psm_q F64 pooled PSM FDR
precursor_q F64 per (peptidoform+charge) FDR

The multi-context q columns and carried peak coordinates are the reasons for the schema bumps; an older reader that assumes the prior column set will misread them. quant.q_filter (see QuantConfig) selects which of these q columns the quant stage filters on.

How it works

Config load path

load_config (main.rs:376-384) reads the --config file to a string and calls Config::from_json (config.rs:1009-1014), which does serde_json::from_str and then validate(). With no --config it returns Config::default() (config.rs:988-1005), which is fully populated from the per-field defaults. --profile <name> then calls apply_profile (main.rs:607, config.rs:1107-1121) on top of the loaded config.

deny_unknown_fields on every section (config.rs:182, 203, 251, ...) means a typo like {"digest":{"min_len":7,"bogus":1}} is a parse error, verified by unknown_key_rejected (config.rs:1146-1149). #[serde(default)] on every section and field means a partial JSON overlays defaults: {"digest":{"min_len":7}} keeps max_len=50, charge_max=3, top_n_peaks=300 (config.rs:1151-1159). The t::<T>() helper (config.rs:163-168) is a terse T::default() used inside the per-struct Default impls.

validate() hard-error checks (config.rs:1019-1100)

Known invalid combinations are rejected at load so they never silently produce a wrong result. This is targeted validation, not proof that every scientifically poor setting is detectable. Defaults always pass.

  1. digest.decoy.strategy == DiannShift -> Invalid: the engine digest produces zero decoys under it, giving an invalid target-decoy FDR (config.rs:1021-1029).
  2. rt_im_train.calibration_method == None -> Invalid: None would silently fall through to the linear fit, so it is rejected and the user must pick linear or loess (config.rs:1030-1036).
  3. extract.retain_top_peaks == 0 -> Invalid: must be >= 1 (1 = legacy single apex) (config.rs:1037-1043).
  4. extract.min_frag_corr not finite or outside [0, 1] -> Invalid (config.rs:1044-1052). 0 disables the gate. Verified by explicit_uncapped_seed_and_invalid_gate_are_distinguished (config.rs:1180-1187).
  5. rescore.folds < 2 -> Invalid: every PSM needs an out-of-fold score (config.rs:1053-1059).
  6. rescore.num_iter == 0 -> Invalid: iterative model training needs at least one iteration (config.rs:1060-1064).
  7. rescore.train_fdr non-finite or outside (0, 1] -> Invalid (config.rs:1065-1072).
  8. Mokapot or NnTorch without rescore.python -> Invalid (config.rs:1073-1082).
  9. classifier = percolator -> Invalid: the adapter is not wired (config.rs:1083-1089).
  10. Entrapment without rescore.entrapment_marker -> Invalid (config.rs:1090-1098).

canonical_json (config.rs:1125-1127) serializes the fully-resolved config; run hashes it with blake3_str into config_hash (run.rs:89) and stores the JSON verbatim in the manifest. There is no separate pretty vs canonical form; it is a plain serde_json::to_string.

apply_profile (config.rs:1107-1121)

Only dia is defined. It sets features.set = Extended, extract.apex_count_window = 5, extract.apex_rt_prior_s = 120.0. Any other name is an Invalid error. All other extraction defaults stay at their conservative baselines; the profile is a convenience shortcut, not a full preset file.

Mass model math (mass.rs, constants.rs)

Neutral mass (mass.rs:66-73): WATER + n_term_mod + c_term_mod plus, per residue, residue_mass(r) + mod_delta. residue_mass (constants.rs:26-51) is the 20-standard-amino-acid table; B, J, O, U, X, Z return None (ambiguous / non-standard) and cause MassError::AmbiguousResidue at parse time. Leucine and isoleucine share 113.084064015 (constants.rs:35-36). is_standard_residue (constants.rs:54-56) is the boolean form (residue_mass(aa).is_some()), used by the digest to reject peptides containing a non-standard residue byte.

m/z conversion (constants.rs:59-62): mass_to_mz(m, z) = (m + z * PROTON) / z. precursor_mz (mass.rs:76-78) is this over the neutral mass.

Fragment generation (mass.rs:82-127): a forward prefix scan builds b ions (b2..b(n-1), dropping b1) and a reverse suffix scan builds y ions (y3..y(n-1), dropping y1 and y2) because b1/y1/y2 are low-information (mass.rs:93, mass.rs:113). Each fragment m/z is (residue_sum_incl_terminus + z * PROTON) / z (mass.rs:97, mass.rs:116). Fragment (mass.rs:45-53) carries ion_type, ordinal, charge, mz, and a stable name (b3, or y5^2 for charge 2, frag_name at mass.rs:130-136). IonType (mass.rs:29-42) is the B/Y enum; IonType::symbol() returns the lowercase 'b'/'y' used in the fragment name. These two series are the only ones the MVP scores (see docs/18_findings_and_decisions.md).

ppm predicates. Three exist and they differ by which mass normalizes the difference; a maintainer must not treat them as interchangeable.

  • ppm_diff(observed, theoretical) (constants.rs:66-68): 1e6 * (observed - theoretical) / theoretical. Signed, normalized by the theoretical mass. Used for reporting a mass error and for ppm_match(obs, theo, tol) = |ppm_diff| <= tol (constants.rs:72-74).
  • ppm_bounds(mz, tol) (constants.rs:78-81): returns (mz - d, mz + d) with d = mz * tol * 1e-6. Symmetric window centered on the query mz, normalized by the query itself.
  • within_ppm(a, b, tol) (constants.rs:92-96): lo = min(a,b), hi = max(a,b), true iff hi - lo <= tol * 1e-6 * lo. Min-relative, symmetric in argument order. This is the canonical fragment-index predicate (fragindex_spec 2.1): it is algebraically hi/lo <= 1 + delta and ln(hi) - ln(lo) <= ln(1 + delta), the last form being what makes log-space binning exact and is proven exact against the log-bin +/-1 probe. within_ppm_three_forms_agree (constants.rs:102-126) checks the three forms round identically away from the edge; within_ppm_edges (constants.rs:128-136) checks edge inclusivity and argument-order symmetry.

The three normalize by theoretical (ppm_diff), by query center (ppm_bounds), and by the smaller mass (within_ppm) respectively, so they disagree at the tolerance edge. Index probing must use within_ppm; do not substitute a ppm_bounds window there.

UniMod subset (mass.rs:13-26): Carbamidomethyl 57.021463735, Oxidation 15.994914620, Acetyl 42.010564684, Phospho 79.966331090, Deamidated 0.984016106, Methyl 14.015650064, Dimethyl 28.031300128, Carbamyl 43.005813726. An unknown name is MassError::UnknownModification (mass.rs:221), never a silent zero.

ProForma-lite parse (parse_peptidoform, mass.rs:143-205): optional N-terminal [Mod]-, residues each optionally followed by [Mod], optional trailing -[Mod]. A [Mod] is a UniMod name or a signed float such as [+15.9949] (parse_bracket, mass.rs:209-225, tries unimod_mass first then f64 parse). Non-alphabetic characters and ambiguous residues are errors. A second mod at the same residue position accumulates (+=, mass.rs:191); the data model has one mods[i] slot per residue plus separate terminal deltas (ParsedPeptidoform, mass.rs:56-62).

Key types and functions

Name file:line What it does
Config config.rs:980 Top-level config; 10 stage sections + mbr + rng_seed
Config::from_json config.rs:1009-1014 Parse (deny unknown) then validate
Config::validate config.rs:1019-1100 Ten hard-error checks
Config::apply_profile config.rs:1107-1121 dia preset only
Config::canonical_json config.rs:1125-1127 Serialize for manifest hashing
residue_mass constants.rs:26-51 20-AA monoisotopic table; None for ambiguous
is_standard_residue constants.rs:54-56 residue_mass(aa).is_some() boolean form
mass_to_mz constants.rs:59-62 Neutral mass -> m/z
ppm_diff / ppm_match constants.rs:66-74 Theoretical-relative signed ppm + tolerance test
ppm_bounds constants.rs:78-81 Query-centered symmetric m/z window
within_ppm constants.rs:92-96 Min-relative canonical index predicate
parse_peptidoform mass.rs:143-205 ProForma-lite parser
ParsedPeptidoform::fragments mass.rs:82-127 b/y ions, drops b1/y1/y2
IonType mass.rs:29-42 B/Y enum; symbol() -> 'b'/'y'
unimod_mass mass.rs:13-26 8-name UniMod subset
Manifest / ArtifactRecord manifest.rs:9-51 Run provenance; new/record/get (manifest.rs:33-51)
record_artifact mumdia-io/src/lib.rs:20-39 Build ArtifactRecord, hash file
RejectionReason rejection.rs:19-113 Ordered identification-loss ladder
Label types.rs:33-51 Target/Decoy; .pin() = +1/-1, .is_decoy()

Core data types (types.rs)

Ion mobility is Option/nullable throughout so one model serves 3D and 4D runs; the MVP is 3D so every IM field is None (types.rs:1-3).

Type file:line Fields / behavior
Peak types.rs:9-13 mz f64, intensity f32, ion_mobility Option<f32> (None for Orbitrap DIA)
IsolationWindow types.rs:17-30 target_mz/lower_mz/upper_mz f64, im_lower/im_upper Option<f32>; covers(mz) is inclusive m/z containment (types.rs:26-29)
Label types.rs:33-51 Target/Decoy (serde snake_case); pin() -> +1/-1 (Percolator), is_decoy()
Ms2Scan types.rs:54-62 scan_index u32, id String, rt_seconds f64, window, m/z-sorted peaks Vec<Peak>

Error types (error.rs)

Both enums derive thiserror::Error; the #[error(...)] string is the Display message. Misconfiguration and bad input fail loudly (see docs/18_findings_and_decisions.md).

Variant file:line Raised when
MassError::Parse(String) error.rs:8-9 peptidoform parse failure: stray -, non-alphabetic char, unclosed [, or no residues (mass.rs:172, 177, 197, 214)
MassError::AmbiguousResidue(char) error.rs:10-11 a residue with no monoisotopic mass (B/J/O/U/X/Z) at parse (mass.rs:183)
MassError::UnknownModification(String) error.rs:12-13 a [...] mod that is neither a UniMod name nor a parseable signed float (mass.rs:221)
ConfigError::Parse(String) error.rs:18-19 serde_json failure inside Config::from_json (config.rs:1011)
ConfigError::Invalid(String) error.rs:20-21 a validate() rejection or unknown --profile (config.rs:1022-1098, 1115)

Candidate-audit reason ladder (rejection.rs)

RejectionReason is the ordered identification-loss ladder written to the candidate-audit table (emit_candidate_audit / mumdia audit). Each candidate's row records the EARLIEST stage at which it was lost, so the aggregate answers "where was each DIA-NN-only precursor first lost?" without conflating later stages. The serialized spelling is SCREAMING_SNAKE_CASE and equals code(); code() (rejection.rs:50-71) is the stable string written to Parquet/JSON without a serde round-trip. stage_order() (rejection.rs:76-97) is the ladder position (0 = earliest); earliest(a, b) (rejection.rs:106-112) keeps the smaller stage_order and Reported never overrides a real rejection; is_rejection() (rejection.rs:100-102) is true for any non-Reported variant.

stage_order Variant / code() Stage lost at
0 PeptideNotGenerated / PEPTIDE_NOT_GENERATED search space (A)
1 ModificationNotAllowed / MODIFICATION_NOT_ALLOWED search space (A)
2 ChargeOutOfRange / CHARGE_OUT_OF_RANGE search space (A)
3 PrecursorMzOutOfRange / PRECURSOR_MZ_OUT_OF_RANGE search space (A)
4 NoValidFragments / NO_VALID_FRAGMENTS search space (A)
5 WrongIsolationWindow / WRONG_ISOLATION_WINDOW search space (A)
6 RtPruned / RT_PRUNED candidate generation (B)
7 CandidateCapReached / CANDIDATE_CAP_REACHED candidate generation (B)
8 NoFragmentTraces / NO_FRAGMENT_TRACES extraction (C, D)
9 NoPeakGroup / NO_PEAK_GROUP extraction (C, D)
10 PeakNotSelected / PEAK_NOT_SELECTED peak/peptide ranking (E)
11 OutcompetedByTarget / OUTCOMPETED_BY_TARGET competition (G)
12 OutcompetedByDecoy / OUTCOMPETED_BY_DECOY competition (G)
13 FailedPrecursorFdr / FAILED_PRECURSOR_FDR FDR + reporting (H)
14 FailedPeptideFdr / FAILED_PEPTIDE_FDR FDR + reporting (H)
15 RemovedDuringReporting / REMOVED_DURING_REPORTING FDR + reporting (H)
255 Reported / REPORTED sentinel: reached the final report (not a loss)

Configuration

Every section is #[serde(default, deny_unknown_fields)]. Below, each field is listed with its default and effect. Fields marked stub or default-off are called out explicitly.

Strategy enums (live variants)

Enum file:line Variants (default in bold) Notes
DecoyStrategy config.rs:17-28 reverse, scramble, diann_shift, none diann_shift rejected by validate; none produces no decoys (invalid FDR)
MatcherKind config.rs:38-42 bucketed, fragindex fragment-matcher backend for search-seed + extract
Enzyme config.rs:46-52 trypsin_p, trypsin cut after K/R, with/without before-P
CalibrationMethod config.rs:56-61 loess, linear, none none rejected by validate (falls through to linear)
FeatureSet config.rs:65-75 minimal (14), rich (44), extended (381) superset battery in stages/features/; count asserted by the feature_sets_sized test (features.rs:1590-1598)
RtPredictorKind config.rs:79-86 native, deeplc DeepLC is a Python sidecar
FragPredictorKind config.rs:90-96 native, ms2pip MS2PIP is a Python sidecar
RescorerKind config.rs:100-123 native_tda, mokapot, nn_torch, percolator, entrapment see RescoreConfig
PeakClaim config.rs:132-157 none, winner_predicted_intensity, proportional, coelution_winner, coelution_proportional, coelution_winner_margin shared-peak apportionment
GateMode config.rs:551-574 apex_pearson, peak_spectral, spectral_entropy, coelution, combined which spectral score min_frag_corr thresholds
UnknownModPolicy config.rs:244-248 error, skip unknown-mod behavior
CompetitionMode config.rs:674-691 winner_take_all, none, features_only, unique_evidence, margin_gated within-group resolution
CompeteGroupBy config.rs:695-702 precursor, apex, peptidoform_charge competition grouping key
RollupMethod config.rs:706-712 top_n_sum, sum protein rollup
PeakWindowMode config.rs:717-729 per_candidate, consensus quant integration window
NormalizeMethod config.rs:737-750 none, median_ratio, median cross-run LFQ normalization; from_token at 754-761
QuantQColumn config.rs:772-789 peptide_q, precursor_q, psm_q, run_psm_q which q column quant filters on
MbrStrategy config.rs:845-855 none, empirical_library, rt_transfer, full stub: only none runs; the rest are config hooks (Stage D3 not wired)
DecoyTransfer config.rs:864-869 permuted_rt, reverse_sequence, both MBR false-transfer null (stub, MBR-only)

Variant semantics (behaviorally-rich enums)

The table names every variant; the ones below carry distinct behavior a maintainer must not confuse. Line refs are config.rs.

MatcherKind (30-42). Fragindex (default) is the log-bin CSR matcher: on narrow-window DIA it is ~1.95x faster in search-seed and ~1.26x in extract with essentially unchanged IDs (HYE B_01 peptides -0.1%). Bucketed is the previous Library::page_search path, retained for A/B comparison and for the AIF full-range-window case, where the min-relative vs query-relative predicate difference shifts IDs more.

RescorerKind (100-123). NativeTda (default): native semi-supervised linear rescorer + target-decoy q, always available. Mokapot (105-106): mokapot Python sidecar over the PIN. NnTorch (107-113): PyTorch semi-supervised MLP sidecar (nn_rescore_worker.py), a nonlinear Percolator/mokapot-style rescorer with CV folds + iterative positive re-selection; on the E.coli benchmark it beats linear mokapot on the same PIN and gains further when the extraction gate is opened; requires rescore.python to point at a torch interpreter. Percolator (114-115): external percolator.exe over the PIN. Entrapment (116-122): treats foreign-proteome PSMs (marked by entrapment_marker) as real negatives, trains a nonlinear GBM sidecar (out-of-fold by base peptide) or a native linear fallback, and reports entrapment-calibrated q-values; the chimeric false matches that in-silico decoys under-model appear as real negatives here.

PeakClaim (125-157) apportions one observed MS2 peak that matches the fragments of several co-isolated, co-eluting candidates (near-universal in wide-window DIA: ~98% of fragment m/z collide within tolerance), to stop a chimeric candidate borrowing a real peptide's peak wholesale. None (default, legacy): every claimant gets the full peak intensity. WinnerPredictedIntensity: winner-take-all by highest predicted intensity for the matching fragment. Proportional: split by predicted intensity. CoelutionWinner (two-pass): a first pass builds each candidate's per-scan summed-matched-intensity elution profile, then the peak goes to the claimant most eluting at that scan (best corroborated by its OTHER fragments), not the one predicting the brightest ion. CoelutionProportional (two-pass): split by per-scan elution-profile height. CoelutionWinnerMargin (two-pass): winner-take-all only when the top eluter's profile height dominates the runner-up by peak_claim_margin, else the peak stays shared (as None).

GateMode (547-574) selects which spectral-agreement score min_frag_corr thresholds, all computed at the gate from data in hand. ApexPearson (default, legacy): Pearson of observed-vs-predicted fragment intensities at the single apex scan (one chimeric scan can dominate). PeakSpectral: Pearson of the peak-integrated observed spectrum (each fragment summed over the elution-peak scans) vs predicted; the standard library-dot-product measure. SpectralEntropy: Li spectral-entropy similarity of the sqrt-transformed apex-scan intensities (spectral_entropy_similarity_sqrt); the full-feature gate search found it the single best gate discriminator (AUC 0.826 / matched-pool recall 69.8% vs apex Pearson 0.781 / 64.5%). Coelution: predicted-intensity-weighted mean co-elution correlation of each matched fragment's XIC to the signature reference over the peak (temporal agreement, orthogonal to intensity). Combined: require BOTH peak-integrated Pearson >= min_frag_corr AND co-elution >= gate_coelution_min.

CompetitionMode (674-691). WinnerTakeAll (default, legacy): keep only the top prelim_score per group. None: keep every candidate (FDR handles ambiguity). FeaturesOnly: same retained set as None, with conflict/contested features carrying the interference signal into rescoring (the name documents intent for the experiment matrix). UniqueEvidence: keep a loser only when unique_fragment_count >= unique_evidence_min_fragments, else winner-take-all fallback (also falls back to WTA when the feature column is absent). MarginGated: remove a loser only when winner_score - loser_score >= margin, else keep it (conservative removal for the low-FDR region). Target/decoy labels stay in the competition key in every mode, so a target never competes against its own decoy.

MbrStrategy (838-855) and DecoyTransfer (857-869) are stubs / config hooks (Stage D3 not wired; all require >= 2 runs). MbrStrategy: None (default) reproduces the chain byte-for-byte; EmpiricalLibrary builds the cross-run consensus anchor library only (M1); RtTransfer adds cross-run expected-RT transfer extraction (M2/M3); Full adds requantification (M5). DecoyTransfer is the MBR false-transfer null (M4): PermutedRt (default) transfers real precursors to a decoupled wrong expected RT; ReverseSequence transfers reverse/scramble decoys at the same expected RT; Both combines them.

CompeteGroupBy (695-702). Precursor (default) groups charge/modification variants of one base peptide, separately within targets and within decoys; Apex additionally buckets by rounded apex RT (apex_rt_tolerance_s); PeptidoformCharge keeps each distinct peptidoform+charge as its own group (precursor-level, as DIA-NN/Spectronaut report), recovering sibling forms the base-peptide grouping collapses. The label stays in the key in every mode: a target never directly competes against its paired decoy in the compete stage. This does not remove target-decoy comparison from FDR: rescore later selects best representatives at each q-value unit and estimates the target-decoy null.

PeakWindowMode (714-729). PerCandidate (default): each candidate's quant window is anchored at the identified apex when available, with a summed-XIC fallback for older scored artifacts; its descent bounds can still be stretched by interference or collapse on sparse peaks. Consensus: the median left/right half-widths of confident peptides applied around each candidate's identified apex. Consensus widths are estimated independently inside each quant invocation, not shared across runs.

QuantQColumn selects which q column quant filters on. PeptideQ (default) uses peptide_q_value. PrecursorQ uses precursor_q and is valid only for a single-run rescore; after pooled rescoring that grouped q is experiment-wide. PsmQ uses pooled q_value. RunPsmQ uses per-source run_psm_q and is the run-local choice after an experiment-wide rescore. Quant has no source selector: slice the scored table by source before pairing it with each run's chromatograms. Changing the q column does not perform that slice.

RollupMethod (706-712): TopNSum (default) sums the top-N most abundant peptides per protein group (top_n_peptides); Sum sums all group peptides. NormalizeMethod (735-761, quant-lfq CLI token, not a config field): MedianRatio (default) is a DESeq-style median-of-ratios size factor over complete-case features, robust to a minority of changing features (does not flatten a spike-in design's real fold changes); Median aligns each run's median intensity (simpler, less robust to composition shifts); None uses raw areas.

DigestConfig (config.rs:181-207)

Field Default Effect
enzyme trypsin_p cleavage rule
missed_cleavages 2 max missed cleavages
min_len 5 min peptide length
max_len 50 max peptide length
decoy.strategy reverse decoy scheme (DecoyConfig, config.rs:170-179)
n_term_met_excision true (bool) when a protein begins with M, also emit the initiator-Met-removed form of its N-terminal peptides, re-checked against min_len/max_len and the standard-residue rule (matches DIA-NN --met-excision); keyed on protein position 0, not any leading M. Omitting it makes the search database structurally miss these excised peptides

PeptidoformsConfig (config.rs:202-231)

Field Default Effect
fixed_mods [{C, Carbamidomethyl}] applied to every matching residue
variable_mods [{M, Oxidation}] optionally applied
max_variable_mods 1 max simultaneous variable mods
charge_min 2 lowest precursor charge
charge_max 3 highest precursor charge
charge_by_basic_residues false when true, ignore charge_min/charge_max and emit charges 1..=1+(#R+#H+#K) per peptide (proton capacity: N-terminus plus one per basic residue)
unknown_modification error error or skip

charge_by_basic_residues restricts the enumerated precursor charge states to what a peptide can physically hold (one proton on the N-terminus plus one per Arg, His, Lys). It replaces the fixed range rather than clamping within it, so a peptide with no basic residue is emitted only at charge 1, and a peptide that cannot reach charge_min is not dropped by charge_min but simply enumerated from 1. Pairs with predict_frag.charge_by_basic_residues (fragments). It changes the search / training / FDR population, so it is benchmark-gated and defaults off. On the imported path, import_diann_lib.py --charge-by-basic-residues applies the same rule (on the reference DIA-NN E. coli library it dropped 7.9% of precursors, almost all charge-3 with a single basic residue).

ResidueMod (config.rs:233-240) is {residue: char, name: String} where name is a UniMod name. The doc comment reserves residue: '*' for "any" and notes terminal mods are handled separately in the MVP (config.rs:236). deny_unknown_fields applies but the struct has no #[serde(default)] (unlike every other section), so both keys are required inside a ResidueMod entry.

PredictFragConfig (config.rs:250-282)

Field Default Effect
predictor native fragment-intensity source (native heuristic or MS2PIP)
rt_predictor native iRT source (native or DeepLC)
charge2_from_precursor_charge 2 add charge-2 fragments for precursor charge >= this
charge_by_basic_residues false when true, keep a b/y fragment at charge z only if z <= 1+(basic residues within the fragment) and z <= precursor charge; supersedes charge2_from_precursor_charge
top_n_fragments 6 fragments kept per candidate
ms2pip_model "HCD" MS2PIP model name
ms2pip_python None interpreter for MS2PIP sidecar
deeplc_python None interpreter for DeepLC sidecar
sidecar_script_dir "scripts" directory holding worker scripts

SearchSeedConfig (config.rs:284-320)

Field Default Effect
fdr_seed 0.01 seed FDR cutoff for calibration anchors
fragment_tol_ppm 20.0 fragment match tolerance
report_psms 5 max PSMs reported per spectrum
min_matched_peaks 4 min matched fragments per seed PSM
top_n_peaks 300 probe only the N most intense MS2 peaks (0 = all)
matcher fragindex matcher backend
two_pass_mass_cal false default-off robust two-pass mass calibration (P3.1)

RtImTrainConfig (config.rs:322-387)

Field Default Effect
calibration_method loess RT calibration (none rejected by validate)
q_train 0.01 q cutoff for calibration anchors
p_rt 0.95 residual percentile for the RT window
rt_window_multiplier 1.0 scales the RT half-window
min_seed_for_calibration 50 min anchors before calibrating
loess_span 0.3 LOESS local-fit fraction
fallback_rt_window_s 120.0 fixed window when calibration cannot fit
finetune_deeplc false default-off DeepLC multitask fine-tune (nondeterministic; needs deeplc_python)
finetune_epochs 25 fine-tune epoch cap (early stopping usually halts earlier)
finetune_patience 10 early-stopping patience
finetune_batch 0 0 = auto-scale batch to seed size
adaptive_rt_window false default-off per-region residual window (P3.2/P3.3)
adaptive_rt_bins 12 RT bins for the adaptive window
rt_window_min_s 1.0 lower clamp on any RT half-window

ExtractConfig (config.rs:389-545)

The largest section; the core extraction stage. Cascade thresholds and apex selection dominate.

Field Default Effect
fixed_scan_window 3 scan-window floor around the apex
frag_tol_ppm 20.0 fragment tolerance
prec_tol_ppm 20.0 precursor tolerance
presence_min_matched 3 tier-(b) min matched fragment count
presence_min_fragments 3 min distinct fragments for acceptance
presence_min_coelution 2 min simultaneously-present fragments over the run
min_frag_corr 0.2 tier-(d) spectral-agreement gate (0 disables; must be in [0,1])
min_matched_fraction 0.0 tier-(c) min fraction of predicted fragments observed
apex_top_fragments 0 signature-fragment apex: sums the observed intensity of the top-K predicted fragments per scan; 0 falls back to a default of 3 (extract.rs:1054-1058), not all-matched
apex_rt_prior_s 0.0 Gaussian RT prior sigma on apex (0 = off)
apex_count_tol 1 fragment-count apex slack
apex_count_window 1 rolling distinct-fragment count width (1 = none; profile dia sets 5)
apex_gaussian_sigma_scans 0.0 Gaussian apex smoother sigma in scans (0 = rolling-sum unchanged; opt-in, benchmark-gated)
emit_window_grid true zero-filled window-grid chromatograms
bucket_size 8192 m/z bucket size (power of two)
peak_claim none shared-peak apportionment mode
emit_contested_features false default-off contested_frac feature (forces two-pass)
peak_claim_margin 2.0 dominance factor for coelution_winner_margin
matcher fragindex matcher backend
min_coelution_run 0 default-off min consecutive co-elution scans
ms1_rescue false default-off MS1-isotope rescue of gate failures
retain_top_peaks 1 K peak groups per candidate (1 = legacy; must be >= 1)
emit_candidate_audit false default-off; in run gates the separate audit stage that writes candidate_audit.parquet (run.rs:405-406); no stage writes <psms>.audit.parquet; no-op for standalone mumdia extract (P0.3)
apex_evidence_rank false default-off evidence-count apex
emit_gate_diagnostics false default-off four gate-score columns
gate_mode apex_pearson which spectral score min_frag_corr thresholds
gate_coelution_min 0.5 second threshold for gate_mode = combined

The min_frag_corr default was relaxed from a historical 0.5 to 0.2 (config.rs:517-522). Every default-off knob is explicitly documented as requiring entrapment/target-decoy FDR validation before use.

FeaturesConfig (config.rs:576-624)

Field Default Effect
set minimal feature set (minimal / rich / extended)
coelution_corr_threshold 0.9 co-elution correlation cutoff
prec_tol_ppm 20.0 precursor tolerance for MS1 features
bound_features true restrict trace features to the elution peak
bound_peak_fraction 1/3 peak-boundary descent fraction of apex height
bound_peak_grace 0 consecutive sub-threshold scans to bridge
bound_from_confident true learn one global peak width from confident seed PSMs
bound_confident_pct 50.0 percentile of confident half-widths as the shared width

CompeteConfig (config.rs:626-664)

Field Default Effect
group_by precursor competition grouping key
apex_rt_tolerance_s 5.0 RT bucket for apex grouping
mode winner_take_all within-group resolution
margin 0.0 score margin for margin_gated
unique_evidence_min_fragments 2 min unique fragments for unique_evidence
emit_competition_audit false default-off writes <out>.compete_audit.parquet

QuantConfig (config.rs:791-836)

Field Default Effect
q_threshold 0.01 cutoff applied to the q column selected by q_filter
top_n_fragments 3 fragments summed per peptidoform
top_n_peptides 3 peptides summed per protein group (TopNSum)
rollup top_n_sum protein rollup method
bound_peak true integrate only over the detected peak window
peak_fraction 1/6 descent threshold for the peak-window walk
peak_grace 1 zig-zag grace (bridge N sub-threshold scans)
peak_window_mode per_candidate per-candidate vs consensus window
reliable_q 0.001 confident-set q for the consensus width
q_filter peptide_q peptide_q, precursor_q (single-run only), psm_q, or run_psm_q
interference_envelope false apex-outward interference envelope on fragment traces before integration (opt-in, benchmark-gated)

RescoreConfig (config.rs:913-965)

Field Default Effect
classifier native_tda rescorer backend
folds 3 CV folds
train_fdr 0.01 positive-selection FDR
num_iter 10 semi-supervised iterations (native)
python None interpreter for a Python rescorer (mokapot/nn_torch/entrapment)
percolator_bin None percolator.exe path
entrapment_marker None protein substring marking spike-in negatives (required for entrapment)
entrapment_exclude None substring that de-marks (own-species)
entrapment_contaminant_markers [] substrings that keep a spike-in hit as a real target
entrapment_ratio 1.0 N_real_lib / N_entrap_lib scaling
strict true fail on a rescorer sidecar failure or unsupported classifier; set false only for explicit compatibility fallback

MbrConfig (config.rs:871-911), stub

Config hooks only; the MBR stage (D3) is not wired into the run chain. Fields: strategy (none default), q_anchor 0.01, min_anchor_runs 2, q_transfer 0.01, rt_window_s 20.0, decoy_transfer permuted_rt, consensus_corr_min 0.0, requant_all false, python None. mbr is the one top-level section with an explicit extra #[serde(default)] attribute (config.rs:985-986). With strategy = none the chain is byte-identical to no MBR.

Top-level Config (config.rs:971-1005)

rng_seed (default 0) plus the ten stage sections and mbr. rng_seed seeds every RNG (decoy scramble, CV fold assignment) for determinism.

Removed / not present

The recent cleanup deleted several dead enums that older docs still mention (DecoySource, ToleranceRegime, ScanWindowMode::PeakWidthDerived, FeatureSet::Custom). They are not in the current config.rs. Do not re-add or document them; if you see them referenced elsewhere, that reference is stale.

Invariants, determinism, gotchas

  • Single source of masses. No crate outside mumdia-core defines a residue mass, PROTON, WATER, AMMONIA, or ISOTOPE_SPACING (constants.rs:10-22). ISOTOPE_SPACING = 1.003354835 is the 13C-12C mass difference used for MS1 envelope spacing. PROTON = 1.007276466812 is the physically correct proton mass, deliberately not DIA-NN's H-atom value 1.007825035 (constants.rs:8-10). All constants are own-derived from CODATA/AME; nothing is copied from another engine (clean-room boundary).
  • Three ppm predicates are not interchangeable. within_ppm (min-relative) is the only one the fragment index may use; ppm_bounds (query-centered) and ppm_diff/ppm_match (theoretical-relative) disagree at the tolerance edge. Substituting one for another shifts which fragments match at the boundary.
  • Validation is targeted and fail-loud for known invalid states. Invalid decoy/calibration settings, impossible rescorer coverage, missing required sidecar settings, and invalid numeric bounds are rejected. It is not a general scientific validator. DecoyStrategy::None is still accepted but produces no decoys and therefore no valid target-decoy FDR; treat it as diagnostic-only.
  • deny_unknown_fields everywhere. A misspelled key fails the whole load, so a config file cannot silently ignore a setting the author intended.
  • Determinism. rng_seed seeds all randomness. canonical_json is a stable serde_json::to_string (field order fixed by struct declaration order), so the same config hashes to the same config_hash across runs. manifest.artifacts and model_identities are BTreeMaps (manifest.rs:28-30), so manifest key order is deterministic. Note finetune_deeplc fine-tuning is itself nondeterministic (no torch seed) and breaks byte-identical reproducibility when enabled.
  • Fragment generation drops b1/y1/y2 unconditionally (mass.rs:93, mass.rs:113); a peptide shorter than 2 residues yields no fragments (mass.rs:85-87).
  • ParsedPeptidoform::neutral_mass uses .expect on residue masses (mass.rs:70), safe only because parse_peptidoform already rejected ambiguous residues. Constructing a ParsedPeptidoform by hand with a non-standard residue byte would panic.
  • Second mod at the same position accumulates (mass.rs:191), which is a behavior difference from engines that drop it; terminal mods are separate fields, not part of mods[i].
  • Recent correctness changes bumped four schemas. PSMS_COMPETED and PEPTIDE_QUANT/PROTEIN_GROUP_QUANT are v2; PSMS_SCORED is v3. Other registry entries remain v1. Readers must honor the registry rather than assume one version globally.
  • content_hash is the file's blake3, not the logical content. Any byte change (compression, column order) changes the hash; it is a change detector, not a canonical-content identity.

How to extend / modify

  • Add a config field. Add it to the section struct, give it a value in that struct's Default impl (every field must have one; #[serde(default)] is at the struct level), and document the effect inline. Do not hardcode a choice the config could express (project convention). If the field can be set to a value that would silently corrupt results, add a validate() check (config.rs:1019-1100) that rejects it with ConfigError::Invalid.
  • Add a strategy variant. Extend the enum, keep #[serde(rename_all = "snake_case")], and (if it is not implementable yet) reject it in validate the way DiannShift and CalibrationMethod::None are, rather than letting it fall through. State default-off status in the doc comment.
  • Add a UniMod modification. Add a name => mass arm to unimod_mass (mass.rs:13-26). Use the PSI-MS/UniMod monoisotopic delta so the Python sidecar adapters map the name. No table is copied from another tool.
  • Add an artifact / bump a schema. Add or bump the constant in schema.rs:7-24. Bump the version whenever the column set changes (as PSMS_SCORED went to 3). Update the producing stage's ArtifactReport (schema_name, schema_version) and its record_artifact call so the manifest and the sidecar report agree.
  • Add a manifest field. Extend Manifest or ArtifactRecord (manifest.rs); both are plain serde structs. Keep new maps as BTreeMap for deterministic key order.
  • Never edit plan.md (gitignored spec, project rule). Keep validated numbers consistent across README, COMPARISON.md, and CLAUDE.md.