Every agent rated against our criteria: problem solved · cadence · time/resources saved · verifiability · build maturity. Every score backed by a real artifact — a PR, a digest, a log. Nothing here is a promise; all of it is a link.
P=Problem · C=Cadence · I=Impact · V=Verifiable · M=Maturity. Total /25. Sources link to the actual artifacts.
| Agent | P | C | I | V | M | Total | Evidence |
|---|---|---|---|---|---|---|---|
| ⚖️ Law-Watch DE · laws |
5 | 4 | 5 | 5 | 4 | 23/25 | Scheduled 06:30. Corpus of 5,936 laws; digest+log+PR pattern proven by sibling agents. ◆ scheduled |
| 🏛️ Bundestag-Watch DE · parliament |
5 | 5 | 5 | 5 | 5 | 25/25 | PR #1: 5 findings, 9 sources, honest "no signal" sections. The gold standard. ◆ ran & verified |
| ✈️ EU Rights-Watch EU · consumer |
5 | 4 | 5 | 4 | 3 | 21/25 | Scheduled 09:00. Protects €4-5B/yr unclaimed EU261. ◆ scheduled |
| 💶 Money-Flow Watch DE · budget |
5 | 4 | 5 | 4 | 3 | 21/25 | Scheduled 10:00. €476.5B budget accountability. ◆ scheduled |
| 🏠 Benefit-Watch DE · benefits |
5 | 5 | 5 | 5 | 5 | 25/25 | PR #1: Wohngeld cut detected with "apply now" citizen tip, +€150m budget line. ◆ ran & verified |
| 🇪🇺 Directive-Watch EU · deadlines |
4 | 4 | 4 | 4 | 3 | 19/25 | Scheduled 11:00. Ends silent late transposition. ◆ scheduled |
| 🏛️⚖️ Court-Watch DE · rulings |
4 | 4 | 4 | 4 | 3 | 19/25 | Scheduled 11:30. BVerfG/EuGH in plain language. ◆ scheduled |
| 🛡️ Abuse-Safety Watch DE · law-change |
5 | 4 | 5 | 4 | 3 | 21/25 | Scheduled 12:00. Keeps SafeVoice's 12-paragraph schema current. ◆ scheduled |
| 📦 Procurement-Watch DE · tenders |
4 | 4 | 4 | 4 | 2 | 18/25 | Scheduled 12:30. Anomaly flags on tenders. Data-source access is the open question. ◆ scheduled |
| 🗳️ Consultation-Watch DE · participation |
5 | 5 | 4 | 5 | 5 | 24/25 | PR #2: 5 windows with deadlines. Participation made actionable. ◆ ran & verified |
| 🎬 Studio Director studio · taste |
5 | 5 | 4 | 5 | 5 | 24/25 | PR #1: 5 line-numbered critiques, caught real hitscan bug, respects the lead's "BOTH" call. ◆ ran & verified |
| 🔧 Gameplay Engineer studio · build |
4 | 5 | 4 | 4 | 2 | 19/25 | Running now (chained to Director's brief). Build-gate enforced. ◆ running |
| 🧪 QA Playtester studio · verify |
4 | 5 | 3 | 5 | 2 | 19/25 | Queued after Engineer. SHIP/HOLD verdict + honest untested list. ◆ queued |
What the scores prove: the loop works — 5 agents ran today, every one produced
cited, logged, human-reviewed output, and the best of them caught a real bug and a real benefit cut.
What they don't prove yet: sustained quality over weeks, real user reach
(the fleet must be found), and the studio's Engineer/QA loop (first run in progress). "Impact" scores are
potential based on public figures — realized impact needs the fleet running daily + people reading it.
The honest label: these are excellent prototypes with real first outputs,
not yet a proven long-term service. That's exactly the right stage to be at on day one.