Michael Ninh
Michael Ninh

AI Engineer

Agentic Systems · AI Operations

I build reliable AI systems that turn messy real-world workflows into tested, observable software.

SpecBuildEvaluateVerifyShipMonitor

Coding agents do much of the implementation. I design the system, architecture, autonomy boundaries, evals and evidence that make the result trustworthy.

Three featured proofs

Different systems. The same engineering discipline.

Choose a capability. Each proof exposes the evidence, the authority boundary and the part that remains deliberately unclaimed.

01The engineering question

Can this AI run a command when execution is supposed to be off?

TrustReady traces explicit execution authority through the software and checks whether deny-by-default enforcement holds before the process sink.

Default stateProcess execution disabled
Explicit authority--allow-exec
VerificationPropagation · sink · fail-closed guard
02Working proof
TRUSTREADY · EXECUTION AUTHORITY

Remove the execution guard. Watch TrustReady refuse GO.

A 30-second interactive replay of a verified ProofWorker path: command checks can reach subprocess.run, but only after explicit --allow-exec authority passes a deny-by-default policy guard.

3real repositories
3architecture families
3/3control suites pass
Fail closedlost guard → NO_GO
Interactive explainer grounded in a verified real-repo scan. Benchmark results are scoped technical assurance, not whole-repository security certification.
How I build

Agents can move fast. The system around them has to stay rigorous.

I design agentic software-development systems with explicit specifications, autonomy boundaries, evals, verification gates and production monitoring.

The goal is not maximum autonomy. It is useful autonomy with evidence: give agents room to execute, give them measurable definitions of done, and keep human judgement where the consequences matter.

01 — SHAPE

Problem first.

Problem → user → constraints → architecture

02 — SPECIFY

Define done.

Requirements → boundaries → acceptance criteria

03 — DELEGATE

Bound autonomy.

Agents execute within explicit autonomy limits

04 — PROVE

Demand evidence.

Tests → evals → benchmarks → adversarial cases

05 — SHIP

Gate release.

CI → deployment gates → production

06 — WATCH

Learn from reality.

Traces → logs → regressions → feedback

Agents build. Evidence earns trust. Humans retain judgement.

When verification fails, the workflow loops backwards. A failure should become a stronger test, boundary, fixture or source rule — not merely another prompt.

A little about me

Product judgement meets hands-on engineering.

Before AI engineering I worked across founder-led e-commerce, product/business development and high-tempo service operations. That background made me unusually interested in the messy bit between a clever model and a workflow people can actually rely on.

I’m looking for an early-career AI engineering / AI operations role where careful reasoning, agentic systems, evidence and real product outcomes matter.

Useful first.
Trustworthy by design.
mikel_ninh@yahoo.de ↗