current project · not public yet
Trust infrastructure for AI agents
Agents take actions. You take responsibility. Observability platforms tell you what happened, after it happened to a real customer. I'm working one step earlier: run the agent in a world you control, high-fidelity simulations of the platforms it operates, and judge it by the side effects it actually causes rather than what it reports. The adversarial states that define a trust boundary almost never show up in live traffic; in a controlled world you can manufacture them on demand.
controlled simulation worlds · deterministic capture & replay · side-effect-level evals · trust envelopes made explicit