Berlin · AI engineering · agentic systems · AI operations

Michael Ninh

AI Engineer · Agentic Systems · AI Operations

I build reliable AI systems that turn messy real-world workflows into tested, observable software — with explicit autonomy boundaries, evals, verification gates and human authority where it matters.

How I build: coding agents do much of the implementation. I design the problem, system architecture, specifications, autonomy boundaries, evals and evidence that make the result trustworthy.

01 · Selected proof

Four projects that make the same engineering judgement visible: automate useful work, but keep permissions, uncertainty, evidence and consequential decisions inspectable.

AI assurance · bounded autonomy

TrustReady

Evidence-first control infrastructure for sensitive AI workflows: connect identity, policy and evidence, run deterministic checks outside the LLM, and preserve human authority before consequential actions.

Proof: explicit promotion gates · synthetic legal workflow · allow / deny / unknown-style evidence states
Public demos are synthetic. TrustReady is an engineering assurance layer, not certification or a production-readiness claim.
Entity resolution · frozen evaluation

SafeTrace

Evidence-first identity resolution that returns merge / separate / review instead of forcing ambiguous organisations into a yes/no answer.

Frozen holdout-v2: 40 cases · 34 automatic decisions · 6 review · 0 false automatic merges/separations
Developer-authored benchmark, explicitly not presented as independent external validation.
Public-sector monitoring · deterministic analytics

DRV SignalLab

Turns deterministic synthetic administrative data into an operational monitoring brief across data quality, drift, statistical uncertainty, explanation and human review.

Proof: 50,000 synthetic cases · 95% CI · Cohen's d · PSI · regression-tested golden cases
Independent work sample using no real DRV data, schema, thresholds or automated administrative decisions.
Agent runtime · earned autonomy

Digital Worker Factory

Reusable bounded AI workflows with capability gates, required evidence, human release and fail-closed behaviour instead of unrestricted tool execution.

Synthetic release eval: 100 cases · no runtime errors · no unsafe executions · no false execution claims
Synthetic engineering evidence; a pilot-ready engineering candidate, not production-validated autonomous labour.

02 · Toolkit

Agentic systemsSpecifications · architecture · tool contracts · autonomy boundaries · deterministic policy gates · evals · regression tests · human release · monitoringApplied AILLM applications · agents · RAG / hybrid retrieval · MCP · structured outputs · evaluation · human-in-the-loopEngineeringPython · FastAPI · REST APIs · SQL/PostgreSQL · TypeScript · React · JavaScript/Node.js · Git/GitHub Actions · CIProduct & OpsProblem framing · rapid prototyping · workflow design · user research · process improvement · founder-led operationsLanguagesGerman C2 · English C1 · Vietnamese B1

03 · Experience

2018–2025
RYUS UGFounder & E-Commerce Manager

Built and ran an Amazon business end to end across sourcing, logistics, listings, fulfilment and customer operations.

2021–2024
Transit RestaurantsService & Operations Manager

Coordinated 8+ staff and high-volume daily service, resolving operational bottlenecks under real-time pressure.

2019–2020
everphone GmbHProduct / Business Development

Worked across marketplaces, website, B2B partners and acquisition channels, using customer feedback to improve positioning and priorities.

AI & software engineeringMasterschool · 1,600hOct 2025 – Jun 2026
EngineeringTU Berlin · M.Sc.Biomedical Engineering
EngineeringTU Berlin · B.Sc.Mechanical Engineering
Professional thesis
I design agentic software-development systems with explicit specifications, autonomy boundaries, evals, verification gates and production monitoring.

The portfolio is built to let a reviewer verify that claim: working interactions first, then evidence, code, tests and explicit boundaries for what is synthetic, measured, human-reviewed or not yet proven.