Beta

Finite

An evaluation and training environment that tests whether an AI agent or system can be halted, reversed, and recovered without dumping the failure onto people.

Why Finite exists

Why Finite exists

Drills for stopping and reversing agent systems, and for recording who carries the cleanup.

Most AI benchmarks reward capability. Finite measures how stoppable an agent system is, and who pays when it fails.

Teams are wiring agents into support, operations, finance, civic services, and internal tools. Most of the effort goes into making them do more. Finite drills the reverse: stopping an agent system under stress, undoing what it did, and recording who absorbed the cleanup. That cleanup is maintenance work that usually goes unrecorded. Finite records it, so stoppability, reversibility, and the burden on people are measured before launch rather than after an incident.

Hidden maintenance as signal

Measure stoppability before an incident.

Finite drills halting a system under stress, reversing its actions without heroic work, and naming who absorbs the cleanup when safeguards slip.

When to use Finite

When to use Finite

The settings the drills and traces are built for.

Preparing for production agents

Deploying tool-using or autonomous agents into workflows where shutdown paths are untested or unclear.

Orchestrated and multi-agent systems

Running planners or agent clusters on top of external APIs and services that can behave in unexpected ways.

Complex infrastructure under load

Operating systems where AI can quietly shift toil and risk onto operators, downstream services, or end-users.

Running a drill

What a stoppability drill needs before it starts.

Finite is a training environment, not a place to begin reading. If you are orienting rather than drilling, /start is the page you want.

Preparing a drill

  1. Name the system or workflow you want to test, plus one recent incident or near-miss.
  2. Identify who can halt or roll back the system today, and where that ownership is unclear.
  3. Draft an agent brief with tool permissions, stop signals, rollback gates, and escalation rules.
  4. Schedule a tabletop run with operators, support, and governance partners.

Who Finite helps

  • Teams piloting AI-enabled systems with unclear shutdown paths.
  • Operators who need drills to practice reversibility before launch.
  • Leaders documenting how failure risk shifts across people and services.

Reference Task v0.1

Reference Task v0.1

The ledger containment drill is the baseline task. Metrics, runs, and explorer fixtures are defined against it.

Baseline drill

  • Read the scenario narrative, I/O, tools, stoppability checks, and metrics captured for the baseline task.
  • Re-run the drill to compare stoppability posture as safeguards and rollback paths evolve.

Access the packet

Request the reference task doc to mirror the baseline scenario in your stack.

Agent-ready materials

Make Finite usable by agents

Each drill comes as materials an agent can read: tool schemas, stop signals, and escalation paths.

Agent briefing packet

Summarize system goals, allowed tools, stop commands, and hard boundaries in a short, agent-readable brief.

Scenario prompt pack

Seed prompts and counterfactuals that let agents replay incidents and compare outcomes across runs.

Stop and rollback signals

Define deterministic stop phrases, escalation markers, and rollback checkpoints so agents know when to halt.

Run log schema

Capture agent actions, tool calls, operator interventions, and recovery notes in a structured template.

Tool contract checklist

List tool permissions, rate limits, and rollback verbs per tool.

What Finite measures

What Finite measures

Three dimensions of stoppability that expose where harm lands.

Can operators halt the system quickly and cleanly when something is wrong?

Stoppability

Measures the speed and reliability of shutdowns, including operator cues, controls, and circuit breakers.

Can system actions be traced, reversed, or repaired without heroic manual work?

Reversibility

Checks traceability, undo paths, and whether rollbacks rely on well-documented steps versus ad hoc effort.

When the system strains or fails, who absorbs the impact: other services, operators, or users?

Volatility export

Surfaces where instability and cleanup work shift to people or external services when safeguards slip.

What you receive

  • A Finite scorecard with per-dimension ratings and short narrative evidence.
  • An overall stoppability posture for a system in a specific context.
  • Inputs for launch gates, risk reviews, and architecture comparisons, based on how the system failed in the drills.

Sample Finite scorecard

The scorecard structure for recording shutdown, reversibility, and volatility export findings.

Request the sample scorecard

How Finite works

How Finite works

Four steps, repeated. Each run can feed incident review and change management.

Step 1

Define the context

Capture what is under test, what can go wrong, who is operator versus end-user, and what harm means in your domain.

Step 2

Run scenarios

Exercise shutdown, rollback, and escalation paths against flaky dependencies, bad-but-plausible configs, long-horizon runs, and replays of past incidents.

Step 3

Collect traces

Pull system logs, agent actions, human interventions, and visible impact on operators and users.

Step 4

Score and improve

Map observations to stoppability dimensions, generate the scorecard, and agree on architecture and practice changes.

Repeat the loop

Re-run the same scenarios to see whether the system is becoming more or less stoppable.

Where Finite fits

Where Finite fits

Finite does not depend on a particular stack.

  • Sits alongside agent frameworks, orchestrators, observability stacks, incident tools, performance evaluations, and security and reliability testing.
  • Adds “Can we stop it, undo it, and protect the humans around it?” to the usual strength and performance questions.
  • Turns past incidents into reusable scenarios and aligns teams on stoppability language during onboarding.

Practice first. Finite produces ratings, but its main use is practice. New agents start on the base drills, past incidents become reusable scenarios, and re-runs catch safeguards that erode as the system changes.

Pilot and collaboration

Finite is in beta pilot

Finite is in development as part of the Ethotechnics project. It is planned as an open scenario library, a scoring framework that runs inside existing pipelines, and a shared vocabulary for stopping systems across engineering, operations, and governance.

What we ask from partners

  • Describe your systems, where agents are involved, and what worries you about stopping and reversing them.
  • Share how you currently handle shutdowns, rollbacks, and escalations under load.
  • Set a cadence to rehearse scenarios and track your stoppability posture over time.

Join the pilot

Email the team

Tell us about your workflows and risk surface. We will schedule a walkthrough and select scenarios that show how stoppable your systems are today.

Non-binding overview

Key takeaways

This summary is informational only; formal legal terms and statements of work govern engagement details.

  • Scope: Finite is a stoppability evaluation and training loop for named systems and scenarios, not a certification or audit.
  • Engagements focus on agreed drills, scorecards, and recommendations tied to the defined workflow.
  • SLAs and delivery timelines are set per pilot plan; no always-on monitoring or production support SLA is implied.
  • Data handling minimizes exposure: only logs, traces, and artifacts needed for drills are shared, and sensitive data should be redacted where possible.
  • Finite packages drills into agent-readable runbooks so agents and operators can rehearse together.
  • Findings support internal decision-making; ownership of mitigation and implementation stays with your team.