What runs underneath the investigation?
A small evidence-first pipeline: preserve the source, resolve identity, build only established relationships, calculate bounded conclusions, keep unknowns visible, then test the whole workflow against a hidden answer and a real browser.
Evidence in. Defensible conclusion out.
The LLM is not the authority. Sources, identity gates, deterministic calculations, explicit uncertainty and human review constrain what can become a conclusion.
The interesting part is where the system refuses to guess.
These are product rules, not prompt suggestions.
A propagated ownership relationship carries source ID and evidence anchor. “Where did this come from?” is answerable.
A nominee shareholder does not become a made-up natural-person UBO. Missing evidence creates a collection gap.
Conflicting DOB, nationality or passport can reject a sanctions near-match despite strong name similarity.
Economic ownership, voting rights and documented non-equity control are calculated and reported separately.
Legal owner, lessee/operator and lender security interest are distinct asset relationships.
Repeated articles tracing back to one origin count as one evidence origin, not five independent confirmations.
Not “looks convincing.” Scored.
The application case has a hidden gold answer and deterministic critical-fail rules. The first-pass runner does not read the gold answer.
Every headline claim has code behind it.
If you are technical, these are the shortest paths into the implementation.
Problem, pipeline, boundaries, live fail-closed case and browser proof.
Core logicOwnership/control enginePath propagation, aggregation, identity/percentage/cycle gates and evidence lineage.
Production boundaryFail-closed analysis contractNo target identity → no graph. Screening handoff remains a research lead, not a sanctions conclusion.
BenchmarkVenatic Analyst Challenge28-record adversarial case, six decisions, 100-point rubric and critical-fail rules.
EvaluationHidden-gold evaluatorDeterministic scoring and critical-failure enforcement.
Blind runInitial analysis runnerReceives the initial case pack without reading the hidden answer.
Research judgementCollection prioritisationRanks unopened sources by expected decision value, authority, novelty and redundancy.
Source criticismSource-independence analysisPrevents syndicated/circular reporting from masquerading as independent corroboration.
RegressionOwnership testsGolden paths plus ambiguous identity, missing percentage, voting/control and cycle guardrails.
Browser QAReal Chromium journeyOpens sources, verifies hidden decisions and runs the research-then-decide interaction end to end.
Benchmark CI95 → 100 gateValidates the case contract, blind score, research-budget score and zero critical failures.
Live boundary CIVenatic ownership boundaryRe-runs the live company case and asserts no shareholder/UBO inference without authoritative evidence.
The live proof is partly about saying “we don’t know.”
The real Venatic public-source pipeline establishes the company identity, but the reviewed sources do not establish an authoritative shareholder list. The correct result is therefore zero ownership edges and zero UBO candidates—not an AI guess.
Acquire allowed public evidence, preserve receipts, resolve established facts, surface gaps, calculate supported paths and prepare research leads for human review.
Invent missing shareholders, turn fuzzy similarity into identity, collapse control concepts, call anomalies criminal conduct, or present a research candidate as a legal conclusion.
Humans steer. Agents accelerate. The product has to pass.
The workflow is outcome-first: define what a defensible investigation must do, encode the failure conditions, let agents implement/review/test, then verify the actual user workflow and change the system when it fails.
Try the investigation, then inspect the implementation.
The demo is intentionally simple. The engineering proof is here for the person who wants to challenge how the answer was produced.