Convene competing scientists
TrueForge opens one persistent case and dispatches bounded specialists with different missions, contexts, and typed outputs.
Independent AI scientists argue, retrieve live evidence, run deterministic experiments, challenge the leader, revise their beliefs, and stop for a human scientist.
A confident first answer is not a scientific verdict.
Defense must make the strongest case. Prosecution must break it. Evidence and experiments decide what survives, not eloquence.
When the leader changes, the old belief, the dissent, and every reason stay visible.
TrueForge opens one persistent case and dispatches bounded specialists with different missions, contexts, and typed outputs.
Bright Data retrieves independent sources and Daytona computes deterministic compatibility, forcing the initial ranking to face new facts.
The old and new rankings remain inspectable, jurors expose disagreement, and promotion halts at the exact scientist approval boundary.
Different missions, different queries, typed outputs, and visible provenance. No agent can quietly grade its own argument.
AdvocateBuilds the strongest evidence-backed case for each candidate without hiding methodology or biological-scope limits.
AdversaryUses independent queries to find contradictions, weak transfer assumptions, and evidence that should lower confidence.
ProvenanceAdmits only source-linked claims with retrieval time, source class, biological scope, methodology flags, and content identity.
RigorChecks whether a paper's design, population, endpoint, and strain specificity justify the weight assigned to it.
ComputationRuns deterministic compatibility and counterfactual calculations in Daytona so the language model cannot invent scores.
ChallengeAttacks the current leader without seeing the desired answer, then sends unresolved objections to jurors and the disagreement analyst.
Candidate A starts at #1 and falls to #3. Candidate B rises from #2 to #1 because admitted evidence plus deterministic computation changes the recorded state.
Defense and Prosecution use different search plans and produce typed arguments for and against every candidate.
The Evidence Clerk accepts only source-linked claims whose biological scope and methodology justify their weight.
The Experimentalist sends deterministic compatibility and counterfactual calculations to an isolated Daytona sandbox.
A blind Red Team and structured jury challenge the leader; the system records the new rank and preserves dissent.
This is the real persistent agent—not a recording. Enter a synthetic R&D objective, watch independent specialists and sponsor tools execute, refresh safely, then approve or reject the exact guarded proposal.
Locating TrueForge…
The control plane will appear when the private runtime responds.
Replay the exported TrueForge event stream: independent Bright Data retrieval, deterministic Daytona execution, adversarial revision, preserved dissent, and the exact scientist approval boundary.
Loading verified run
Defense and Prosecution retrieve independently. Every admitted claim keeps its URL, retrieval time, source class, biological scope, methodology flags, and content hash.

Compatibility, sensitivity, and counterfactual checks run as deterministic code in an isolated sandbox. Inputs, outputs, and hashes stay attached to the decision.

Read-only research can run autonomously. Promotion cannot. The scientist sees the immutable proposal, evidence, dissent, rollback scope, and SHA-256 hash before any write. Drag to inspect it.
Sources, experiments, candidate revisions, and dissent remain attached to the verdict.
An adversarial multi-agent microbiome R&D workflow. Independent specialists argue, retrieve evidence, run deterministic experiments, revise candidate rankings, preserve dissent, and stop for a scientist.
No. The public case is synthetic, experimental, and non-clinical. It does not diagnose, prescribe, recommend treatment, or claim clinical validation.
TrueForge owns the persistent case, specialist delegation, MCP tool calls, sandbox events, context continuation, and the exact human approval boundary.
Bright Data retrieves query-specific sources with provenance and recovery metadata. Daytona executes deterministic compatibility and counterfactual code. Both change the recorded ranking rather than decorating the interface.
Averages can hide meaningful scientific disagreement. The system stores each verdict and uses a disagreement analyst to classify where the jurors diverge and why.