TraceCore, the Deterministic Episode Runtime, prioritizes deterministic core stability, auditability, and adoption scaffolding before optional attestations and ecosystem extras.
- Deterministic first: stable runner contracts (CLI + artifact schema), frozen task manifests, reproducible baselines.
- Auditability with restraint: integrity hashing now; signatures/attestations once schemas are stable.
- Adoption-focused: CI-ready templates, minimal-start examples, small deterministic task library with clear budgets.
- Scope discipline: optionalize heavy ledger/blockchain and certifications; gate multi-agent/async behind proven single-agent determinism.
Deliverables
- Freeze runner contracts (CLI + artifact schema) and land deterministic baseline export/compare with shared local/CI TOML.
- Ship IO audit diffs in Trace Viewer plus taxonomy regression tests for validator outcomes.
- Enforce artifact integrity via hashed bundles and publish GuardedEnv + validator normalization security review. Exit criteria
- Reference agents can replay frozen tasks reproducibly (local + CI) with zero schema drift.
- Validator taxonomy events are emitted deterministically in regression suites.
- Integrity hashing is on by default with documented verifier steps.
Deliverables
- Expand deterministic task catalog with frozen manifests, CI policy templates, and “minimal start” examples.
- Ship focused adapters for priority stacks (LangChain in v0.4; OpenAI/Anthropic deferred to v0.5 once LangChain usage hardens) with deterministic shims and budget enforcement.
- Produce structured trace exports (e.g., OTLP) plus an episode config schema for swapping models/tools under budgets. Exit criteria
- Teams can adopt TraceCore via turnkey templates that cover pass/fail gates, artifact diffing, and budget alerts.
- At least three external agents run on the expanded task catalog without contract tweaks.
- OTLP/episode config exports flow into a sample monitoring pipeline without manual patching.
Ground this phase in today’s agent_bench baseline --compare / diff_runs flow so roadmap promises map to the existing deterministic diff surface area.
Deliverables
- Formalize frozen task/version policy, evidence bundles, and contributor playbook.
- Enable optional signing/attestation (e.g., Cosign) once schemas are stable; keep blockchain/IPFS storage opt-in.
- Deliver Trace diff CLI (
tracecore diff run1 run2) and richer failure taxonomy UX by:- Surfacing a dedicated CLI entrypoint that loads run artifacts, honors the existing structured diff schema, and can emit JSON + OTLP-compatible exports for monitoring pipelines.
- Rendering the richer taxonomy panel (failure type + termination reason) by default so validator outcomes remain visible without extra flags.
- Budgeting <10s turnaround on baseline hardware for common diff sizes (≤1k steps) and documenting the runbook for investigating slower cases.
- Updating docs/tutorials so teams know how to extend taxonomy metadata and route the structured diff into dashboards. Exit criteria
- Evidence bundle format is versioned, documented, and consumed by at least one pilot integrator, with Trace diff CLI output linking to the same bundle metadata.
- Signing/attestation passes smoke tests for deterministically hashed bundles without blocking unsigned flows, and CLI diff tooling can verify whether compared runs were signed.
- Trace diff CLI highlights regression deltas and taxonomy shifts in <10s for baseline scenarios on reference hardware, including OTLP/JSON exports that downstream monitors ingest without manual patching.
Phase 4 (3–4 quarters): Scale and readiness for v1.0 (Status: complete — shipped in tracecore v1.0.0)
Deliverables
- Performance: parallel episode execution under bounded resources plus resource/budget monitoring.
- Reliability: red-team tool-call standardization, hardened regression suites, and steady minor release cadence toward v1.0.
- Metrics: CI pilot adoption dashboards, reproducibility pass rates, time-to-diagnose regressions instrumentation. Exit criteria
- Parallel runs on bounded hardware show ≤5% nondeterminism rate with back-pressure controls.
- Nightly regression packs cover all frozen tasks with <1% flake rate.
- Metrics dashboards show upward trends for CI adoption and declining MTTR for regressions across two consecutive releases. Shipped
tracecore run batch—ProcessPoolExecutorworker pool,--workers,--timeout,--strict-spec, P50/P95 wall-clock aggregation.runner/isolation.py— realmultiprocessing.spawnprocess isolation; replaces 5-line stub.wall_clock_elapsed_s— required artifact field (excluded fromartifact_hash); spec v1.0 normative.runner/metrics.py—compute_metrics,compute_all_metrics,compute_mttr.tracecore runs metrics/tracecore runs mttrCLI commands +GET /api/metrics+/metricsdashboard.tracecoreconsole-script entry point;tracecore versioncommand.spec/tracecore-spec-v1.0.md+spec/artifact-schema-v1.0.json— promoted from v0.1.test_action_contracts.py— action contract regression suite across all registered tasks.- Dashboard Run button fix (async executor);
__init__.pyagent dropdown fix.
Status recap: Phase 4 delivered trace diff CLI, trust pipeline (signing/verification), OTLP exports, and taxonomy UX. Phase 5 builds on that foundation to expand task variety, harden runtime architecture, and operationalize diagnostics so TraceCore can support production benchmarking.
Deliverables
- Task portfolio expansion: Scenario packs (security triage, customer support escalation, autonomous ops), multi-agent orchestration harness, and updated SPEC governance to support ≥3 multi-agent tasks.
- Observability & diagnostics: Provider-agnostic LLM telemetry module, replay diff CLI/dashboard UX, ledger usage guide, and MTTR playbooks tying telemetry to troubleshooting.
- Architecture & runtime evolution: Safe timeout manager (subprocess/async), optional reasoning hooks, distributed runner alpha with artifact streaming, scheduler controls.
- Ecosystem acceleration: Production-ready LangChain/OpenAI/OpenClaw guides, leaderboard ingestion design doc + preview endpoints, plugin discovery UX.
- Documentation & UX refresh: FAQ rewrite, debugging playbook, contributor onboarding guides, OpenClaw tutorial refresh, CLI help improvements.
- Testing & migration tooling: Expanded negative suites, schema migration tool + CI hook, hosted LLM integration tests, distributed-runner nightly acceptance job.
- Performance & scalability: Load/stress harness (≥1k episodes) with perf dashboards, artifact compression/streaming options, regression alert thresholds.
Exit criteria
- All P0 checklist items closed with documentation and regression tests.
- Telemetry/replay tooling demonstrably reduces MTTR to <15 minutes in pilot feedback.
- Distributed runner executes ≥10 concurrent tasks without budget violations or orphaned threads.
- External contributor publishes a signed plugin/task using Phase 6 docs/tooling.
- Checklist shows ≥80% completion of P1 scope with no blocked P0 items; deferred work carries rationale.
- P0 focus: contract freeze + deterministic compare flow remains top priority.
- Signing/attestation: optional after schema stability; not mandatory for baseline use.
- Framework/provider priority: start with LangChain and OpenAI/Anthropic APIs before expanding to others (e.g., CrewAI) as demand warrants.
| Priority | Focus | Milestones | Exit criteria |
|---|---|---|---|
| P0 (Critical) | Deterministic core + baseline hygiene | Lock runner contracts (CLI + artifact schema), release deterministic baseline compare flow, ship shared local/CI TOML config | Reference tasks run reproducibly across local and CI; schema-breaking changes require explicit version bump |
| P1 (High) | Adoption scaffolding | Expand deterministic task catalog, publish CI policy templates, improve trace and failure analysis UX | Teams can adopt a standard gating workflow with artifact diffs and clear failure taxonomy |
| P2 (Medium) | Trust + ecosystem scale | Formalize frozen task/version policy, improve plugin/registry contribution path, document trust/repro evidence model | External contributors can add tasks/plugins under stable contracts; release-to-release comparability is auditable |
| Risk area | Why it matters | Mitigation |
|---|---|---|
| Contract churn in early APIs | Breaks adoption and invalidates historical comparisons | Introduce versioned contracts, deprecation windows, and schema alerts in CI |
| Task growth without quality bar | More tasks can reduce signal if determinism slips | Require deterministic validator checks, frozen manifests, and taxonomy regression gates |
| CI integration friction | Teams may skip adoption if setup is heavy | Provide opinionated templates, minimal-start examples, and turnkey artifact diff scripts |
| Analysis UX lag | Artifact volume can outpace debugging usefulness | Prioritize top failure modes, guarantee tracecore diff <10s on frozen baselines, and document taxonomy/evidence schema mapping |
- Adoption: count of CI pilots running TraceCore nightly/weekly; goal is ≥5 before v0.9.
- Determinism: reproducibility pass rate across frozen tasks; target ≥99% with automated alarm on drift.
- Budget discipline: median tool-call budget consumption vs. ceiling per task; provide dashboard slices per release.
- Time-to-diagnose regressions: track mean time from failure detection to root cause using Trace diff tooling; target <1 day by Phase 4.
- Evidence/attestation readiness: measure percentage of bundles shipped with integrity hashes and optional signatures once Phase 3 unlocks.