An evaluation and training environment that tests whether an AI agent or system can be halted, reversed, and recovered without dumping the failure onto people.
Drills for stopping and reversing agent systems, and for recording who carries the cleanup.
Most AI benchmarks reward capability. Finite measures how stoppable an agent system is, and who pays when it fails.
Teams are wiring agents into support, operations, finance, civic services, and internal tools. Most of the effort goes into making them do more. Finite drills the reverse: stopping an agent system under stress, undoing what it did, and recording who absorbed the cleanup. That cleanup is maintenance work that usually goes unrecorded. Finite records it, so stoppability, reversibility, and the burden on people are measured before launch rather than after an incident.
Hidden maintenance as signal
Measure stoppability before an incident.
Finite drills halting a system under stress, reversing its actions
without heroic work, and naming who absorbs the cleanup when
safeguards slip.
Four steps, repeated. Each run can feed incident review and change management.
Step 1
Define the context
Capture what is under test, what can go wrong, who is operator versus end-user, and what harm means in your domain.
Step 2
Run scenarios
Exercise shutdown, rollback, and escalation paths against flaky dependencies, bad-but-plausible configs, long-horizon runs, and replays of past incidents.
Step 3
Collect traces
Pull system logs, agent actions, human interventions, and visible impact on operators and users.
Step 4
Score and improve
Map observations to stoppability dimensions, generate the scorecard, and agree on architecture and practice changes.
Repeat the loop
Re-run the same scenarios to see whether the system is becoming more or less stoppable.
Sits alongside agent frameworks, orchestrators, observability stacks, incident tools, performance evaluations, and security and reliability testing.
Adds “Can we stop it, undo it, and protect the humans around it?” to the usual strength and performance questions.
Turns past incidents into reusable scenarios and aligns teams on stoppability language during onboarding.
Practice first.
Finite produces ratings, but its main use is practice. New agents start on the base drills, past incidents become reusable scenarios, and re-runs catch safeguards that erode as the system changes.
Finite is in development as part of the Ethotechnics project. It is planned as an open scenario library, a scoring framework that runs inside existing pipelines, and a shared vocabulary for stopping systems across engineering, operations, and governance.
What we ask from partners
Describe your systems, where agents are involved, and what worries you about stopping and reversing them.
Share how you currently handle shutdowns, rollbacks, and escalations under load.
Set a cadence to rehearse scenarios and track your stoppability posture over time.
Join the pilot
Email the team
Tell us about your workflows and risk surface. We will schedule a walkthrough and select scenarios that show how stoppable your systems are today.