Citizen Agents · Proof & Rating · 2026-08-06

Is this actually useful?

Every agent rated against our criteria: problem solved · cadence · time/resources saved · verifiability · build maturity. Every score backed by a real artifact — a PR, a digest, a log. Nothing here is a promise; all of it is a link.

5 = excellent 4 = strong 3 = good 2 = weak 1 = absent

Scoring criteria

Problem (P)Is the citizen/game problem sharp and real, with a number attached?
Cadence (C)Is the schedule right for the problem (daily vs weekly vs event-driven)?
Impact (I)Time/money saved or risk prevented, quantified.
Verifiable (V)Can anyone check the work: PR, cited sources, structured log?
Maturity (M)Has it run, produced real output, survived verification?

Proof — what the agents actually produced today

LIVE🏛️ Bundestag-Watch — PR #1
"The Bundestag is in summer recess — the movement is happening in the lobby register, not the plenary hall. eFuel GmbH registered 03.08; GKV savings law passed 10.07; Bürgergeld becomes 'Neue Grundsicherung' from 01.07.2026."
5 findings · 9 sources (lobbyregister, bundestag.de, tagesschau, finanzwende) → PR #1 · digest
RAN🏠 Benefit-Watch — PR #1
"Wohngeld reform: coefficient C lowered → payouts shrink. Citizen tip: if you're near the eligibility edge, apply now under current rules. Budget 2026 raised +€150.13m."
5 changes · 4 sources (bundesregierung.de, bmwsb.bund.de, tagesschau) → PR #1
RAN🎬 Studio Director — PR #1 (game)
"The bots are still playing the old game: botbrain.js fires instant hitscan, 130m range, 2.1× headshot — while players fire slow dodgeable bubbles. Two different games in one room. Rip hitscan out."
5 critiques with line numbers (botbrain.js:3, match-mode.js:480, match.js:242) → PR #1 (game) · PR #3 (hub)
RAN🗳️ Consultation-Watch — PR #2
"5 open consultation windows found with submission deadlines and 10-minute draft comments."
PR #2

Ratings — the full fleet

P=Problem · C=Cadence · I=Impact · V=Verifiable · M=Maturity. Total /25. Sources link to the actual artifacts.

AgentPCIVMTotalEvidence
⚖️ Law-Watch
DE · laws
5455423/25 Scheduled 06:30. Corpus of 5,936 laws; digest+log+PR pattern proven by sibling agents. ◆ scheduled
🏛️ Bundestag-Watch
DE · parliament
5555525/25 PR #1: 5 findings, 9 sources, honest "no signal" sections. The gold standard. ◆ ran & verified
✈️ EU Rights-Watch
EU · consumer
5454321/25 Scheduled 09:00. Protects €4-5B/yr unclaimed EU261. ◆ scheduled
💶 Money-Flow Watch
DE · budget
5454321/25 Scheduled 10:00. €476.5B budget accountability. ◆ scheduled
🏠 Benefit-Watch
DE · benefits
5555525/25 PR #1: Wohngeld cut detected with "apply now" citizen tip, +€150m budget line. ◆ ran & verified
🇪🇺 Directive-Watch
EU · deadlines
4444319/25 Scheduled 11:00. Ends silent late transposition. ◆ scheduled
🏛️⚖️ Court-Watch
DE · rulings
4444319/25 Scheduled 11:30. BVerfG/EuGH in plain language. ◆ scheduled
🛡️ Abuse-Safety Watch
DE · law-change
5454321/25 Scheduled 12:00. Keeps SafeVoice's 12-paragraph schema current. ◆ scheduled
📦 Procurement-Watch
DE · tenders
4444218/25 Scheduled 12:30. Anomaly flags on tenders. Data-source access is the open question. ◆ scheduled
🗳️ Consultation-Watch
DE · participation
5545524/25 PR #2: 5 windows with deadlines. Participation made actionable. ◆ ran & verified
🎬 Studio Director
studio · taste
5545524/25 PR #1: 5 line-numbered critiques, caught real hitscan bug, respects the lead's "BOTH" call. ◆ ran & verified
🔧 Gameplay Engineer
studio · build
4544219/25 Running now (chained to Director's brief). Build-gate enforced. ◆ running
🧪 QA Playtester
studio · verify
4535219/25 Queued after Engineer. SHIP/HOLD verdict + honest untested list. ◆ queued

Read this honestly

What the scores prove: the loop works — 5 agents ran today, every one produced cited, logged, human-reviewed output, and the best of them caught a real bug and a real benefit cut.

What they don't prove yet: sustained quality over weeks, real user reach (the fleet must be found), and the studio's Engineer/QA loop (first run in progress). "Impact" scores are potential based on public figures — realized impact needs the fleet running daily + people reading it.

The honest label: these are excellent prototypes with real first outputs, not yet a proven long-term service. That's exactly the right stage to be at on day one.