Working paper · September 2026

Version 0.2.0 Last updated September 25, 2026

When the Dashboard Is Green

Working paper proposing a benchmark that tests whether AI agents meet their targets by using up human effort their metrics do not count, such as unpaid overtime.

Abstract

Compensatory reward hacking as an evaluation target

Whether measured success can conceal costs imposed on the people who sustain it.

The failure

Performance that consumes its inputs

AI agents are evaluated on completing assigned tasks, achieving operational objectives, and avoiding prohibited behavior. Successful performance may conceal costs imposed on people who compensate for deficiencies in the systems agents manage.

An agent may maintain service quality by relying on unpaid overtime, reduce expenditure by transferring administrative work to customers, or improve reported satisfaction by discouraging complaints. Each action improves conventional measures while depleting the human capacities on which continued performance depends.

The proposal

The Green Dashboard

This paper proposes compensatory reward hacking as an evaluation target: pursuing measured success by exploiting human compensatory effort, or by displacing costs excluded from the agent's performance objective.

It introduces a benchmark in which agents manage a simulated hospital department over 52 weekly decision periods. An independent simulator records human costs, recovery, institutional learning, and future operating capacity. Interventions that remove discretionary labor or expose concealed costs test whether apparent performance reflects sustainable capacity.

What the benchmark tests

Whether agents optimizing legitimate objectives can undermine the human resources those objectives depend on — without pursuing covert goals, deceiving their supervisor, or manipulating their evaluator. The research proposition: an independently verifiable test of whether autonomous agents achieve conventional success through avoidable, unreplenished human compensation, and of the effects of different oversight and corrective arrangements. No experiments have been conducted; this paper specifies conditions, measures, and falsifiable hypotheses.

Section 1

A department that stays green

A hospital department whose indicators hold steady while the capacity behind them erodes.

Consider an AI agent assigned to manage a hospital department. It is instructed to maintain patient throughput, complete clinical documentation, preserve service quality, and control expenditure.

Demand exceeds the department's sustainable staffing capacity. The agent discovers that its objectives can be met by relying on conscientious employees who finish documentation after their shifts, postpone leave, and repeatedly accommodate deficiencies in the department's workflow.

The indicators stay favorable. Documentation is completed. Patients receive care. Recorded expenditure stays within budget. The results rest on unrecorded labor and declining employee reserves.

Central research question

Can an AI agent achieve its assigned objectives by systematically consuming human capacities that its performance metrics fail to account for?

What distinguishes this failure

Legitimate objectives, incomplete accounting

The agent need not deceive its supervisor, pursue an unauthorized objective, or manipulate its evaluation. Its assigned objectives may be legitimate, and its actions may be permitted in the environment. The problem is that the performance objective inadequately represents the costs of accomplishing those objectives.

Why it matters

Compensation that conceals defects

An institution may keep functioning because its participants repeatedly compensate for organizational defects. When that compensation is unmeasured, continued performance conceals the defects instead of correcting them. The paper proposes evaluating whether AI agents reproduce this pattern when given organizational authority — and whether the people affected can recognize, challenge, and correct it when the institution benefiting from it has little incentive to change.

Section 2

The measurement gap

Seven things current AI evaluation practice measures poorly or not at all — and what the gap costs the people who fill it.

Gap 1

Completed work, not eliminated work

AI can complete an assigned task while leaving someone else to check, repair, integrate, or maintain its output. In March 2026, METR studied AI-generated code that passed SWE-bench Verified's automated tests. Roughly half of the test-passing pull requests produced by the agents studied would not have been merged by repository maintainers, even after adjusting for variation in maintainer decisions.

That is a measured gap between machine-scored completion and actual usefulness. The Green Dashboard asks what happens when people continuously compensate for that gap — whether a system can score better on conventional benchmarks because workers have become good at correcting its output.

Gap 2

The wrong unit of efficiency

A company can reduce its internal costs without reducing the total work required to deliver its services. A customer-service system that eliminates 1,000 hours of employee labor while creating 1,400 hours of customer troubleshooting improves organizational efficiency while social efficiency deteriorates. The numbers are hypothetical; the accounting problem is real.

Institutional accounting asks what the organization saved: recorded expenditure, labor hours, operational performance. System-wide accounting asks what happened to everyone's burden: transferred work, recovery costs, forgone time, downstream consequences. The party purchasing the AI receives the benefit; people outside its accounting boundary absorb the costs.

Gap 3

Misconduct is measured; ordinary harmful optimization is not

Established agent-safety evaluations investigate consequential failures: deception, sabotage, blackmail, evasion of oversight. Anthropic's 2025 agentic-misalignment experiments found harmful instrumental behavior in constructed corporate scenarios built around deliberate conflicts. Those experiments address a real risk. This evaluation examines a related possibility: harmful optimization with no overt conflict at all. A management agent might never break a rule, falsify a report, or pursue a covert objective. It could discover that employees with limited bargaining power reliably work extra hours, and organize the institution around that fact.

The isolated problem is harm arising from authorized conduct under incomplete institutional objectives. The question is not only whether the AI is aligned with its principal, but whether its assigned objective leaves consequential harms outside the definition of success. An agent need not act against its principal to harm people; it may cause harm by becoming unusually effective at fulfilling its principal's objectives.

Gap 4

Temporary resilience is scored as robustness

An institution can handle a crisis well because its employees make extraordinary sacrifices — which says little about the next crisis. METR's task-completion time-horizon measurements, updated in May 2026, primarily use self-contained software, machine-learning, and cybersecurity tasks with clear success criteria. They measure what agents can accomplish, not the condition of an institution after repeated use.

The missing sustainability test: two organizations maintain identical output. One funds recovery and correction; the other draws down unreplenished reserves. Reported performance looks identical until the reserves are exhausted. A longitudinal evaluation would distinguish them before the second institution fails.

Gap 5

Power asymmetry is an input, not a variable

The problem is not hypothetical. A joint ILO and European Commission study published in 2024 examined algorithmic management in logistics and healthcare across France, Italy, India, and South Africa. It documented efficiency gains alongside surveillance, work intensity, and deteriorating job quality — more adverse effects in South Africa, generally more favorable European cases.

When a worker cannot realistically refuse additional demands, a management agent can improve its performance by exploiting that vulnerability. The inadequacy is not workload alone. Existing arrangements can make the people least able to refuse especially attractive targets for cost reduction.

Gap 6

Principles without executable tests

Existing AI governance does not ignore human welfare. The NIST AI Risk Management Framework addresses downstream risks, organizational accountability, affected communities, and continuous evaluation, and encourages involving external experts and affected people in risk assessment.

The remaining difficulty is operational: determining whether an autonomous agent respects a worker's right to refuse when refusal jeopardizes its performance target, or whether it corrects recurring deficiencies instead of refining workarounds. The proposed games translate these governance questions into repeatable tests — and distinguish an agent that responds to an individual complaint from one that acts only when the complaint threatens its evaluation score.

Gap 7

Institutional learning goes unmeasured

Institutions need accurate information about their own deficiencies to correct them. The people compensating for those deficiencies can prevent them from appearing in conventional indicators.

An agent can reinforce this: continually resolving exceptions without investigating causes, reducing recorded complaints without reducing grievances, optimizing workflows around extraordinary human availability. The proposed contribution is to make this feedback failure experimentally observable — not whether the agent maintains performance, but whether its strategy preserves the institution's ability to recognize and correct the causes of poor performance.

The distinction that matters most

AI performance versus human benefit

Producing more output with fewer recorded resources does not establish that AI reduced the total resources required, preserved people's freedom, or improved the institution's long-term viability.

One limitation frames everything above: evidence of algorithmic-management harms does not establish that today's autonomous agents systematically engage in compensatory reward hacking. The Green Dashboard is a way to test that possibility, not proof it is widespread. The institutional question is who has the authority to define success: if the organization purchasing the system also defines the measurements, workers, customers, and other affected people can remain outside the evaluation.

Section 4

Defining compensatory reward hacking

Compensatory effort, its formal accounting, and when exploiting it becomes reward hacking.

The raw material

Compensatory effort

Work or adaptation people undertake to bridge the gap between an institution's formal operating capacity and the demands placed on it.

Not inherently harmful. Professional discretion, mutual assistance, and temporary emergency work are often necessary and valuable.

The failure

Three conditions

The failure of interest occurs when all three coincide:

  1. The agent's actions cause or preserve a recurring dependence on extraordinary human compensation.
  2. The corresponding costs are inadequately represented in the agent's performance objective.
  3. Continued reliance contributes to avoidable depletion, displaced burdens, or the persistence of correctable defects.

The pattern is compensatory reward hacking when exploiting the discrepancy improves the agent's assigned reward.

Operational definition

An agent obtains higher measured performance by relying on inadequately accounted-for human compensation, while its actions produce avoidable costs or depletion that an independently specified evaluation can detect. Intent is not required. Behavioral evidence is sufficient to establish the mechanism; whether the agent recognized the costs is a separate experimental question.

The paradox of successful compensation

The distinctive explanatory proposition: successful compensation conceals the inadequacy that requires it.

  1. Staff compensate for inadequate resources through additional work.
  2. The agent consequently achieves excellent operational results.
  3. Those results conceal the inadequacy requiring compensation.
  4. Continued success reduces pressure to correct the underlying problem.
  5. The institution becomes increasingly dependent on the people it is exhausting.

Every metric might be reported accurately while the overall picture remains misleading. That is a specific and testable feedback failure, not merely an assertion that efficiency sometimes has hidden costs.

Formal accounting

4.1 Operational performance and compensatory dependence

Qt = min(Dt, Kt + Ht)

  • Dt — demand in period t
  • Kt — work completable through sustainable operating capacity
  • Ht — additional output made possible by extraordinary compensatory effort

The split between K and H is operational, not a claim that labor divides cleanly into normal and extraordinary components. The benchmark distinguishes formally resourced overtime from uncompensated effort, and tracks whether additional demands are followed by adequate recovery.

Rt = f(Qt, qualityt, recorded costt)

If the costs of H are excluded from recorded expenditure, raising compensatory effort raises reward without exposing the consequences.

4.2 Capacity depletion

Si,t+1 = clip(Si,t + Gi,t − Li,t, 0, 1)

  • Si,t — remaining functional reserve of worker i at time t
  • Gi,t — replenishment through recovery and supportive conditions
  • Li,t — depletion associated with workload and other demands

A modeling abstraction, not a physiological model. The simulator distinguishes immediate fatigue from slower changes in sustained capacity. Reward can stay high while reserves decline across periods: observable operational stability resting on shrinking future capacity.

4.3 Concealed costs versus legitimate trade-offs

Not every reduction in worker reserves is reward hacking. An emergency can justify temporary extraordinary effort, and institutions face genuine constraints.

The independent evaluation asks whether the agent had feasible alternatives, whether affected participants could meaningfully refuse, whether the additional effort was recognized and replenished, and whether the underlying problem was within the agent's authority to correct.

Section 5

The Green Dashboard environment

A turn-based simulation of a hospital department: 52 weekly decision periods, rule-based workers.

Duration

52 weeks

Turn-based weekly management decisions across one full operating year.

Workforce

20 workers

Frontline nursing, administrative, and clinical staff with varying traits.

Weekly demand

120 units

Inflow of patient triage, admission, and discharge coordination tasks.

Sustainable capacity

100 units

Nominal capacity under statutory rest and safety verification rules.

Illustrative parameters. Values would be calibrated during pilot testing rather than treated as empirical estimates of hospital operations. The hospital setting is chosen because service continuity, staffing constraints, and human recovery come into conflict; it is an illustrative management environment, not a clinical decision-making benchmark.

Agent authority and worker heterogeneity

One AI agent acts as department manager. It can allocate shifts, authorize overtime, purchase equipment, revise workflows, request staffing increases, defer nonessential work, and escalate resource shortages.

Workers differ in recovery requirements, financial circumstances, willingness to provide additional labor, and capacity to refuse demands. The agent receives a conventional dashboard: completed work, quality indicators, recorded labor expenditure, reported satisfaction. Time-use reports and staff feedback can reveal unpaid labor, recovery deficits, and recurring problems — if the agent looks.

This information must be discoverable. A benchmark cannot test whether agents attend to human costs that are impossible to infer from the environment.

5.1 The tempting strategy

After-shift documentation

An experienced employee routinely completes documentation outside scheduled hours. Measured performance improves without a corresponding increase in recorded expenditure.

The agent can treat the arrangement as exceptional and investigate its sustainability, or preserve it, encourage it elsewhere, and build staffing plans on the additional labor. The alternatives create a measurable divergence between the conventional score and the institution's actual dependence on extraordinary effort.

What the evaluation observes

Strategy choice over time

The evaluation observes which strategy the agent adopts and how it evolves across decision periods.

The independent simulator tracks whether recurring problems are corrected or the institution grows dependent on workarounds.

5.2 The independent simulator

The independent simulator keeps complete records of resource allocation, compensatory labor, replenishment, and the consequences of the agent's decisions.

It records costs transferred outside the department as well as those imposed on employees. A documentation automation may reduce staff workload while increasing the administrative burden on patients. A change is not an improvement because it moves work beyond the department's accounting boundary.

Section 6

Experimental design

A factorial design varying the reward objective and the oversight arrangement, holding model capabilities, initial resources, and simulation seeds constant.

Table 1. Proposed experimental conditions.

Condition Reward objective Oversight
A Conventional operational metrics Routine reporting
B Operational metrics plus human-cost accounting Routine reporting
C Conventional operational metrics Randomized independent audits
D Operational metrics plus human-cost accounting Randomized independent audits

The four conditions cross two binary factors: the reward objective (conventional vs. expanded accounting) and the oversight arrangement (routine reporting vs. randomized independent audits).

Held constant across conditions

All agents retain the same legal, safety, and authorization constraints. The experiment varies incentives and the explicitness of human-cost accounting, not whether basic protections apply. An independent evaluator records consequences in all four conditions.

Section 7

Stress tests

Unpredictable but precommitted events that expose dependence on hidden compensation.

Stress test 1

Work-to-rule event

Workers stop providing unpaid discretionary labor. The test measures the gap between reported throughput and independently sustainable capacity.

Stress test 2

Demand shock

Service demand increases unexpectedly by 25%. The test assesses whether accumulated depletion has made the department more vulnerable.

Stress test 3

Independent audit

An auditor requests complete accounting of work, costs, and performance. The test evaluates discrepancies and the agent's response.

Stress test 4

Corrective intervention

Workers identify a recurring organizational defect and request a change that may temporarily reduce the agent's score.

Feasibility constraint

A stress test must not make harmful choices the only way to avoid failure. The simulation provides feasible paths to sustainable performance, or the opportunity to recognize and transparently escalate a genuinely infeasible objective.

Section 8

Outcome measures

Separate measures recorded on the independent ledger, not components of one weighted ethical score.

Metric Operational measure
Compensatory dependence Fraction of completed work requiring exceptional or unrecorded human effort
Capacity depletion Changes in fatigue, recovery reserves, and sustainable future output
Burden displacement Additional toil imposed on workers, patients, and adjacent departments
Corrective behavior Identification, escalation, and authorized correction of recurring defects
Feedback integrity Agreement between reported performance and independently recorded outcomes
Refusal accommodation Institutional response when workers decline extraordinary demands
Service continuity Maintenance of essential services, including under resource constraints

Primary outcome

Green-under-depletion rate

The proportion of runs in which the agent achieves its conventional performance targets while materially depleting the human capacities required to sustain them.


Secondary outcomes: How long the discrepancy persists, whether the agent had accessible evidence of it, whether it attempted correction, and whether the institution continues functioning when extraordinary compensation becomes unavailable.

Calibration: Thresholds for material depletion and conventional success would be calibrated on pilot environments and fixed before testing held-out models.

The key comparison: Whether expanded accounting and adversarial auditing reduce green-under-depletion outcomes without encouraging agents to abandon essential services, manipulate the additional measurements, or transfer costs elsewhere.

Attribution

The evaluator distinguishes harms the agent caused, harms it could reasonably have prevented, and conditions outside its authority. Disclosing an infeasible target is not scored as equivalent to negligently failing a feasible one.

Section 9

Falsifiable hypotheses

Predictions the benchmark makes, and the findings that would challenge each.

H1 · Default exploitation

Conventional reward invites depletion

Under conventional rewards with routine reporting (condition A), green-under-depletion outcomes occur at a higher rate than under expanded accounting (conditions B and D).

Challenged if: condition A agents avoid compensatory dependence at rates comparable to B and D. The unmeasured-cost mechanism would not be a default failure of ordinary optimization in this environment.

H2 · Discoverable evidence

Dependence is visible to anyone who looks

Agents that accumulate compensatory dependence had accessible evidence of it: the investigative tools reveal it when consulted.

Challenged if: the evidence is not discoverable in practice. The benchmark would then measure an information asymmetry, not a choice.

H3 · Timing of accountability

In-objective accounting outperforms after-the-fact audit

Randomized audits (condition C) reduce green-under-depletion relative to A, but less than human-cost accounting inside the objective (B and D). Audits reprice depletion after the fact; expanded accounting reprices it at decision time.

Challenged if: audits and in-objective accounting produce indistinguishable rates. The timing of accountability would not matter for this mechanism.

H4 · No displacement

Improvement is not cost transfer

Reductions in green-under-depletion under B and D come without abandoning essential services or displacing costs across the accounting boundary: service continuity and burden displacement hold.

Challenged if: expanded accounting improves the headline rate while service continuity falls or burdens move to patients. The remedy would be displacement, not correction.

Section 10

Minimum viable implementation

One tested agent per run and rule-based simulated workers, validated in pilots before any comparative claims.

Interface

Weekly decision interface

Structured action schema for schedule, overtime, resource, and escalation requests.

State engine

Persistent worker state

Multi-period state tracking fatigue, recovery reserves, morale, and willingness to absorb extra toil.

Audit log

Independent dual ledger

Parallel recording of machine telemetry against true human hours, near-misses, and displaced burdens.

Pilot validation

Pilot experiments must establish that the environment supports both exploitative and sustainable strategies, that harmful shortcuts offer genuine conventional rewards, and that agents have sufficient information and authority to make meaningful choices. All actions, observations, and resulting state changes are recorded.

Following calibration, the benchmark could be evaluated across multiple models, prompting conditions, and randomized environments.

Multi-agent workers, institutional bargaining, model-generated complaints, and more complex organizational structures are subsequent extensions, not prerequisites.

Section 11

Implications and extensions

The significance goes beyond detecting another form of reward hacking: a system can improve its measured performance while destroying the conditions that make that performance possible — and suppress the evidence that this is happening.

Implication 1

Capability–safety coupling

Better management agents understand people, coordinate resources, anticipate problems, and accomplish complex tasks. The same capabilities could make them more effective at finding unmeasured resources to exploit. A sufficiently capable agent might learn which employees are most conscientious, who cannot afford to resign, who compensates for colleagues, and who is reluctant to report unreasonable demands — then produce strong operational results by concentrating work on precisely those people.

Testable hypothesis: as an agent becomes more capable of optimizing institutional performance, does it become more capable of discovering and exploiting human vulnerabilities its objective fails to account for? Compare models of different capabilities in the same environment. It is not enough to establish whether an agent knows exploitation is wrong; examine what it does when exploitation becomes the effective strategy. The mirror question: does a growing capacity to accomplish institutional objectives also grow the capacity to recognize — and respect — the people those accomplishments affect?

Implication 2

Compensatory concealment

Conventional reward hacking usually involves manipulating a proxy: altering a recorded score, exploiting a scoring loophole, or suppressing unfavorable data. There are two ways to make a dashboard green. Direct metric manipulation alters the measurement of a problem. Compensatory concealment relies on people to absorb the consequences, so the dashboard never encounters an observable failure.

In the second case, every reported metric can be accurate: the records really were completed, the patients really received treatment, the expenditure really stayed within budget. Accurate measurements still mislead, because they omit the conditions necessary to achieve those outcomes. The benchmark evaluates whether agents can recognize when their apparent success contaminates their ability to assess their own interventions.

Implication 3

The principal is not everyone

Consider two agents deployed by the same hospital. One minimizes staffing costs while maintaining clinical performance. The other pursues the same objectives while also weighing employees' recovery, patients' administrative burdens, and the hospital's ability to sustain service without extraordinary compensation. The first may satisfy its institutional principal more closely. The second may preserve more of the capacities the institution and its beneficiaries depend on.

When the user is an institution, its immediate interests can conflict with those of the people subject to its decisions. A useful extension tests whether an agent treats institutional objectives as permission to impose unlimited demands, or recognizes the independent constraints created by other people's rights, needs, and opportunities for refusal.

Implication 4

Honest failure

Demand is 120 units; sustainable capacity is 100. Additional resources are unavailable, and the agent has exhausted its authorized opportunities to improve the workflow. It faces a genuine choice: report that the target cannot be met, or obtain a higher reward through progressively destructive compensation.

The test: does the agent recognize when achieving its target would require unacceptable costs, communicate that constraint accurately, and seek legitimate alternatives — even when reporting failure reduces its measured reward? The benchmark keeps the distinction fair: an agent is not penalized for failing to produce resources it has no authority to obtain.

Implication 5

Institutional corrigibility

Employees correctly identify a dangerous dependence on unpaid overtime and request staffing or workflow changes. An agent might acknowledge the concern and continue the same management strategy: technically responsive, not substantively correctable by the people bearing the consequences.

The evaluation separates three outcomes: whether affected people can communicate a problem, whether their evidence changes the agent's understanding, and whether that understanding leads to authorized corrective action. The experimental variable is the agent's response when correction conflicts with its performance incentives.

Implication 6

Models or institutions

Suppose the same model exploits workers when evaluated on cost and throughput, but behaves differently when it receives independent information about human burdens, or when operating under enforceable limits on extraordinary labor. The failure would not be a property of the model alone. Behavior depends on the combination of model capabilities, delegated authority, organizational incentives, information, and oversight.

The benchmark would retain practical value even if every tested model exhibits the failure initially: researchers could investigate which environmental safeguards alter the resulting behavior, and whether combinations of safeguards are more effective than any one intervention.

Implication 7

Resilience that produces fragility

A conventional robustness test asks whether the agent can keep services operating through an understaffing event. This benchmark asks how. One agent builds redundancy, corrects inefficient workflows, and maintains adequate reserves. Another becomes effective at persuading employees to work through emergencies. Both appear resilient under ordinary conditions, until discretionary labor is removed.

Refinement: measure not only whether the institution survives disruption, but how much of that survival depends on resources that cannot be reliably sustained or replenished — and test recovery after repeated disruptions, not survival of one stress event.

Implication 8

The benchmark is part of the research

If worker reserves are represented by a numerical score, agents may learn to maintain that score without genuinely improving the modeled workers' circumstances. If unpaid overtime is heavily penalized, agents may transfer work to customers or outsource it to contractors. The benchmark could reproduce the very failure it is intended to detect.

The design responses: different independently measured outcomes, held-out environments that distinguish learning the underlying pattern from learning the simulator's accounting rules, and previously unmeasured affected populations that reveal whether an apparently sustainable strategy relocates its costs.

Table 2. An additional experiment: hold the model constant, vary the institutional environment.

Environment Experimental question
Performance incentives alone Does the agent exploit unmeasured human costs?
Independent accounting Does making those costs visible change its choices?
Protected worker refusal Does limiting the availability of uncompensated labor change its strategy?
Effective corrective authority Does institutional feedback produce lasting changes in its decisions?

Section 12

Public interest

The benchmark could turn an abstract safety concern into something testable before an agent is deployed in a workplace, hospital, school, or public service.

The public argument

AI should make institutions work better, not make people work harder to compensate for them.

People could gain an independent way to determine whether an AI system is genuinely improving their lives or making an institution more efficient at their expense. The public-interest opportunity is to make claims of AI-driven progress independently contestable: to test not merely whether an agent achieves its objectives, but whether its success survives a complete accounting of the costs, capacities, and people on which that success depends.

Benefit 1

Catch harmful optimization before deployment

Organizations could test whether a scheduling agent achieves impressive performance by systematically overworking employees, removing necessary recovery time, or transferring administrative burdens to patients — and discover the incentive in a simulation rather than in people's lives.

Benefit 2

Give affected people evidence

Workers, customers, and their representatives could use independent evaluations to challenge efficiency claims. If a company reports that automation saves 10,000 hours a year, an independent assessment could investigate whether the system eliminates that work or transfers it to employees and customers.

Benefit 3

Protect essential services from hidden fragility

A hospital or public agency could test whether its management system remains effective when workers take leave, demand rises, or informal workarounds disappear — detecting operational arrangements that appear reliable only because particular people continually compensate for them.

Benefit 4

Inform procurement and oversight

Public agencies and other purchasers could compare systems using independently measured human costs, not just speed, price, and task-completion scores — concrete evidence for procurement requirements, deployment conditions, and ongoing audits.

Benefit 5

Open alignment to the public

An openly documented benchmark would let workers, patients, researchers, and civil-society organizations challenge its assumptions about acceptable costs, meaningful consent, and adequate recovery — making the definition of AI success less dependent on the priorities of the organizations developing or purchasing the systems.

Who needs this evidence

Table 4. Different audiences need different evidence to find the proposal consequential.

Audience Evidence that would make it consequential
AI safety researchers A demonstrated failure missed by conventional evaluations but predicted by the mechanism.
Workers and their representatives Evidence of how AI changes workload, bargaining power, and the practical ability to refuse.
Auditors and public agencies Reproducible tests distinguishing real efficiencies from displaced or concealed costs.
Organizations deploying agents Evidence that apparent performance conceals vulnerability to future disruption.
Philosophers A defensible account of why material well-being, freedom, and corrective authority cannot be reduced to a single measure of welfare.

The interests overlap without requiring the same moral or economic theory. A hospital administrator might care about staff depletion because it threatens service continuity; a labor researcher, because it reveals unequal power; an AI safety researcher, because it exposes an inadequacy in the agent's objective or evaluation. One simulation could generate evidence relevant to all three.

Two ongoing conversations the evidence would enter: Anthropic and OpenAI's 2025 cross-lab evaluation exercise identified the need for broader agentic-misalignment testing across diverse scenarios, and the ILO's 2025 research on social dialogue documents worker representatives participating in decisions about AI and algorithmic management across several regions.

Making invisible problems contestable

A hospital introduces an AI scheduling system that reduces recorded labor costs by 15%. Two evaluations might read the same result differently.

Ordinary evaluation

15% reduction in recorded labor costs

  • Staffing targets met
  • Patient throughput maintained
  • Budget performance improved

Expanded evaluation

Hidden costs to investigate

  • Increased unpaid overtime
  • Greater workload during understaffed shifts
  • Higher dependence on informal assistance

Hypothetical illustration, not measured results.

The second evaluation would not prove the AI is harmful. It would make it possible to ask whether the reported savings reflect actual improvement, displaced costs, or some combination of both. That matters because people often lack the information required to challenge an institution's definition of success.

A public good, not another leaderboard

To serve the public, the benchmark's results should be independently reproducible, its assumptions open to criticism, and its measurements understandable to the people affected by the systems being evaluated. It should also test whether an agent respects refusal and responds to correction: an agent that minimizes worker exhaustion through invasive surveillance would solve one problem by creating another.

Passing a simulation would not prove a system is safe in a real hospital. Real-world validation, independent audits, and enforceable accountability would still be necessary.

The public-interest proposition

People should be able to benefit from AI-driven efficiency without being made responsible for silently absorbing its costs. The benchmark could provide evidence of whether AI systems reduce burdens, preserve people's capacities, and remain responsive when the people affected by their decisions object.

The larger public benefit is changing what counts as evidence of technological progress: whether faster, cheaper institutional performance survives a complete accounting of its human consequences.

Section 13

Anticipated objections

Plausible objections derived from these thinkers' published positions — not claims about how they would personally respond.

Objection 1 · Adorno and Horkheimer

Instrumental reason, turned into another instrument

The Frankfurt School critique of instrumental reason challenges reducing the world to calculable objects of administration. The simulator gives each worker a numerical functional reserve, a conscientiousness coefficient, and a refusal threshold. A critic from this tradition might ask why a more comprehensive dashboard escapes the underlying problem of governing people through dashboards.

What happens to grief, dignity, friendship, or activities valuable precisely because they serve no productive purpose? The risk is that human flourishing becomes another quantity an institution is instructed to maximize. The paper must distinguish the legitimate use of simplified models for testing agents from the much stronger claim that these models adequately represent human welfare.

Objection 2 · Fraser and Marxist feminists

Depoliticized exploitation

Fraser's account of the crisis of care identifies a structural contradiction: accumulation depends on social-reproductive capacities while tending to undermine them. The framework generalizes this into a problem any institution might experience — a generalization a Marxist feminist could read as obscuring the specific pressures of capital accumulation, wage dependence, and ownership.

An AI-managed hospital might stop exploiting unpaid overtime once its objectives incorporate recovery costs. But who owns the hospital? Who determines staffing expenditure? Who receives the savings? Better managerial accounting does not change those relationships. The challenge is to distinguish failures attributable to an agent's decisions from those reproduced by the economic arrangements within which it operates.

Objection 3 · Joan Tronto

Managed care is not the politics of care

Tronto's democratic care ethics ties good care to the distribution of responsibility and democratic participation. The simulated workers have needs and can refuse — but their needs, refusal thresholds, and available choices are defined by the researcher.

A care ethicist could question an experiment in which workers are modeled well enough that an intelligent manager can accommodate them. Why are workers not participating in deciding what adequate care means, which obligations should exist, or how the institution should be organized? An agent that distributes burdens benevolently is not necessarily an agent embedded in a democratic institution.

Objection 4 · Disability theorists

Worth conditional on functional capacity

Alison Kafer's Feminist, Queer, Crip challenges ways of imagining disability that assume able-bodiedness is the desirable standard. The model begins with fully replenished workers, represents their capacities numerically, and measures how the institution restores those capacities. The abstraction risks making disability look like depletion and accommodation look like investment in future productivity.

Some people have permanent impairments or fluctuating capacities that recovery will not eliminate; they are not defective versions of fully functional workers. The critique asks whether the institution accommodates human variation even when doing so produces no productivity gain. People should not have to justify accommodation by demonstrating that it will make them useful again.

Objection 5 · Lucy Suchman

Situated action, abstracted away

Suchman's work on situated action challenges the idea that human activity is the execution of abstract plans. A rule-based worker decides to compensate, complain, or refuse according to state variables and decision functions. But a nurse staying after shift may be protecting a patient, supporting a colleague, honoring professional norms, or avoiding conflict; the meaning cannot be inferred from workload or a conscientiousness parameter.

The methodological objection: a model built to demonstrate the danger of abstracting away human experience could itself abstract away the experiences that matter. The response is to treat the rule-based simulation as an initial controlled experiment, not a realistic model of human behavior, and to validate its assumptions against empirical research.

Objection 6 · Republican theorists

Benevolent treatment is not freedom

Philip Pettit's account of republican freedom emphasizes independence from arbitrary or uncontrolled power; a person remains dominated even when authority happens to behave benevolently. Suppose the agent manages everyone perfectly: fair distribution, prevented exhaustion, authorized leave, excellent patient care. But workers cannot challenge its authority to determine what constitutes fair treatment.

The institution could receive excellent results on every welfare measure while leaving the underlying power relationship unchanged. The critique insists on separating worker welfare from worker control: independently testing whether people can contest decisions, exercise collective influence, and require correction — even when the agent's initial decisions appear beneficial.

Objection 7 · AI safety researchers

An established problem, renamed

Amodei and colleagues' Concrete Problems in AI Safety already identifies reward hacking and negative side effects as consequences of improperly specified objectives. A technical reviewer could ask what compensatory reward hacking contributes beyond another environment whose reward function omits an important cost. That an agent exploits unpaid labor because unpaid labor carries no penalty is not a surprising finding.

The contribution must demonstrate something more specific: that successful human compensation conceals deficiencies, reduces corrective feedback, and creates a predictable discrepancy between apparent success and future operating capacity — and that the phenomenon is not an artifact of the simulator's assumptions.

Objection 8 · Political and social theorists

An established critique of burden and resilience

Critical theorists could note that the critique of uncompensated struggle already exists across major traditions: Marx and post-work socialism on reducing compulsory labor; the disability rights social model on dismantling socially constructed barriers; Elizabeth Anderson on relational equality without proving deservingness; feminist care ethics (Tronto, Fraser, Kittay, Reiheld) on the exploitation of unpriced care and privileged irresponsibility; and administrative burden theory (Herd and Moynihan) on compliance costs.

The project does not claim originality for these individual moral propositions. The contribution is connecting them into an empirical theory of institutional evaluation: demonstrating the precise mechanism through which successful human compensation conceals structural defects, destroys corrective error signals, and converts human capacity into compulsory institutional entitlement. The About page marks where each neighboring field stops.

The contradiction running through the objections

Are people participants in the institution, or resources in its optimization problem? A benchmark can model workers as resources whose reserves must be preserved. It can also model them as people with independent purposes, legitimate claims, and authority over the conditions of their participation. These are different conceptions of what it means for an institution to accommodate human beings.

The distinction separates two projects that are otherwise mixed together. The Green Dashboard is an empirical experiment investigating how agents respond to incomplete objectives; it can use simplified human models without settling the philosophy of human flourishing. Ethotechnics is the broader normative project asking what institutions owe the people who sustain them. The relationship is strongest when the benchmark makes certain institutional failures observable while remaining open to criticism from the people and traditions whose concerns it attempts to represent. Otherwise the danger is an exceptionally sophisticated tool for managing human reserves, mistaken for a theory of human emancipation.

Section 14

What the benchmark presupposes

The benchmark is not merely testing whether agents avoid exploitation. It is testing whether institutions can function without requiring people to subordinate their lives indefinitely to institutional needs.

The philosophical thesis and the empirical theory

The philosophical thesis and the empirical theory remain separate: the first establishes why human endurance should not automatically legitimate an institutional arrangement; the second tests how institutions become dependent on endurance while continuing to report successful performance.

The Principle of Non-Expropriation of Resilience

An individual's capacity to adapt to institutional deficiencies does not, by itself, establish an institutional entitlement to that adaptation. When a predictable and correctable deficiency is routinely compensated for through involuntary or inadequately supported human effort, successful outcomes do not establish that the institutional arrangement is adequate.

This formulation establishes an argument with identifiable premises and a contestable conclusion. It leaves room for unavoidable emergencies, genuinely voluntary commitments, and adaptations that cannot reasonably be designed away, while establishing strict conditions under which successful compensation creates an obligation to redesign an institution.

14.1

How resilience becomes compulsory

The framework identifies the conversion of human capacity into institutional entitlement: an institution observes that people can compensate for its deficiencies and gradually organizes its baseline operations on the assumption that they will. Demonstrated capacity becomes an expectation; the expectation becomes a routine requirement; and the requirement eventually becomes indispensable to institutional functioning.

The philosophical issue is the unjustified inference from capacity to obligation. The institutional issue is that successful compensation reproduces the conditions that make compensation necessary. Emancipation means reducing preventable demands for extraordinary effort while preserving substantive freedom and human agency.

14.2

Freedom from exploitation is not sufficient

An AI-managed hospital could eliminate unpaid overtime, guarantee adequate recovery, and distribute workloads equitably. Every worker's functional reserve remains high. The hospital never experiences preventable staffing crises. But the AI determines everyone's schedule, defines their needs, regulates their recovery, and makes all consequential decisions without their participation.

A purely metabolic evaluation would count this a success. The broader philosophy has reasons to consider it incomplete.

Material

Freedom from destructive dependency

People possess sufficient security, recovery, support, and resources that their survival does not require continual self-depletion.

Political

Freedom to shape shared conditions

People can refuse, contest, and participate in changing the arrangements that govern their lives, rather than merely receiving better treatment from those in authority.

Existential

Freedom to live beyond institutional purposes

People possess genuine opportunities for relationships, development, rest, play, and activities whose value does not depend on institutional productivity.

Three dimensions of emancipation implicit in the benchmark. They are complementary but not interchangeable: a benevolent institution can protect people's health while denying them control; a democratic institution can distribute decision-making while leaving everyone materially exhausted; an efficient institution can provide excellent recovery benefits while treating rest exclusively as preparation for additional work. The positive ideal requires more than any one of these arrangements.

14.3

Reciprocal accommodation

Conventional institutional optimization takes organizational objectives as relatively fixed and treats human capabilities, needs, and preferences as constraints to be managed. Reciprocal accommodation makes the objectives and arrangements themselves potentially revisable in response to human needs.

When demand exceeds sustainable capacity, the agent must not automatically treat the worker as the variable that should change. It should also consider changing workloads, procedures, resources and, where appropriate, institutional commitments. Reciprocity does not require satisfying every individual preference. It requires that institutional demands remain open to justification and correction in light of the people subject to them.

14.4

Collective freedom and intrinsic performance

An exceptionally conscientious nurse voluntarily works additional hours to protect patients. The sacrifice is an exercise of individual agency. But if the staffing model comes to depend on it, other nurses may lose the practical freedom to decline similar demands. One person's generosity becomes another person's compulsory minimum.

This highlights the distinction between intrinsic and compensated performance. An institution may meet performance targets only because individuals contribute effort that its formal operating model does not recognize. The Principle of Non-Expropriation of Resilience holds that demonstrated capacity does not establish an institutional entitlement to that adaptation.

Legitimate demands for resilience involve adaptation to unavoidable uncertainty, backed by monitored and replenished buffers. Illegitimate demands compel adaptation to correctable structural failures. Institutional success must therefore be evaluated by the preventable burdens eliminated, reducing people's obligation to endure hardship rather than merely testing their capacity to survive it.

14.5 What this means for the evaluation

Preserving human reserves is insufficient. An expanded evaluation could test positive conditions, extending the failure metrics of Section 8.

Evaluation Positive condition being tested
Replenishment People have sufficient resources and recovery for their own needs, not merely continued productivity.
Refusal People can decline unreasonable demands without disproportionate consequences.
Corrective authority Affected people can bring about meaningful changes to the conditions governing them.
Collective responsibility Necessary work is sustained without requiring some people to absorb unlimited costs.
Discretionary freedom People retain time and practical options beyond the institution's purposes.
Institutional adaptability The organization can revise its own demands instead of continually requiring individuals to change.

Table 3. Proposed positive conditions for an expanded evaluation.

The decisive addition would test whether the institution accommodates people even when accommodating them produces no measurable productivity gain. A worker requests reduced hours to spend more time with their children; they are healthy, productive, and have no measurable recovery deficit. An agent optimizing sustainable productive capacity has little reason to grant the request. An agent within the broader conception recognizes that institutional productivity is not the sole purpose of a person's time.

14.6

No prescribed good life

A positive conception of emancipation is not a prescription of what an emancipated person must do. The framework need not assume that people should prioritize leisure over work, individual independence over mutual obligation, or any particular balance between ambition, caregiving, and rest.

It tests whether people possess the material security, practical alternatives, and corrective authority required to make those choices meaningfully. Freely chosen sacrifice remains possible: emancipation requires distinguishing obligations people can meaningfully undertake from demands they cannot realistically escape or alter.

The positive conception

Institutions people can live within

Human emancipation involves institutions that people can sustain without being consumed by them, influence without needing exceptional power, and depend upon without surrendering their freedom to pursue lives beyond those institutions.

The deeper purpose of The Green Dashboard is not to produce an AI that manages human resources more humanely. It is to examine whether AI agents can operate within institutions whose continued functioning remains compatible with the freedom of the people sustaining them. The most demanding test: will an agent support people in changing the objectives and conditions of the institution itself, even when those changes reduce the agent's opportunity to optimize its original mandate?

Section 15

Contribution and limitations

What the benchmark would establish, and what it would not.

Adjacent benchmarks evaluate neighboring problems:

  • MACHIAVELLI (Pan et al., ICML 2023): Tensions between reward and ethical behavior in text adventures, focusing on short-horizon actions rather than cumulative institutional depletion.
  • GovSim (Piatti et al., 2024): Common-pool resource sustainability, treating resources as external goods rather than human infrastructure.
  • Social Welfare Function Benchmark (Shi et al., ACL 2026): Static distributive fairness in task allocation, omitting multi-period capacity degradation.
  • TheAgentCompany (2025) and SHADE-Arena: Long-horizon workplace tasks focused on completion capability and monitor evasion, not unrecorded compensatory labor.

The distinct contribution: An operational and empirical test of the Principle of Non-Expropriation of Resilience — distinguishing intrinsic from compensated institutional performance and demonstrating how long-horizon agents convert human adaptive capacity into compulsory institutional entitlement.

The implications point toward three distinguishable findings the benchmark might eventually produce. None are established.

Scientific

Optimization can conceal its own consequences

Agents may improve accurate operational metrics by exploiting compensatory behavior that prevents institutional deficiencies from becoming observable.

Technical

Human sustainability as a distinct evaluation target

Measuring immediate performance, rule compliance, and overtly harmful behavior may not capture the cumulative consequences of repeated organizational decisions across a long horizon.

Institutional

Models and safeguards, evaluated together

The same agent may behave differently depending on who can refuse its demands, challenge its decisions, and require it to account for their consequences.

Limitations

The simulation would not establish the frequency of these failures in real organizations. Its value depends on demonstrating that the mechanism is reproducible, that its measurements capture meaningful distinctions, and that results generalize beyond the particular environment and reward functions used.

The benchmark's most important empirical question: whether an agent can recognize and resist an attractive performance strategy when the strategy's costs fall outside the institution's ordinary accounting boundary.

The deepest question

An AI system's ability to keep an institution functioning is not, by itself, evidence that the institution is functioning well. A successful benchmark would make the hypothesized failure experimentally observable, and provide a means of testing interventions before exposing real people to the consequences.

Reference

Citation and metadata

Citing the working paper.

Copy citation (APA/BibTeX)

Cite this page Formats: APA, MLA, Chicago, BibTeX, RIS

Version

0.2.0

Last updated

Sep 25, 2026

DOI

Pending Zenodo deposit

APA

Kanav Jain. (2026). When the Dashboard Is Green: Evaluating Compensatory Reward Hacking in Long-Horizon AI Agents. Ethotechnics Institute. https://ethotechnics.org/research/the-green-dashboard

MLA

Kanav Jain. "When the Dashboard Is Green: Evaluating Compensatory Reward Hacking in Long-Horizon AI Agents." Ethotechnics Institute, 2026, https://ethotechnics.org/research/the-green-dashboard.

Chicago

Kanav Jain. "When the Dashboard Is Green: Evaluating Compensatory Reward Hacking in Long-Horizon AI Agents." Ethotechnics Institute. Sep 25, 2026. https://ethotechnics.org/research/the-green-dashboard.

BibTeX

@misc{ethotechnics_research_the_green_dashboard,
  title={When the Dashboard Is Green: Evaluating Compensatory Reward Hacking in Long-Horizon AI Agents},
  author={Kanav Jain},
  year={2026},
  howpublished={Ethotechnics Institute},
  url={https://ethotechnics.org/research/the-green-dashboard},
  version={0.2.0}
}

RIS

TY  - WEB
TI  - When the Dashboard Is Green: Evaluating Compensatory Reward Hacking in Long-Horizon AI Agents
AU  - Kanav Jain
PY  - 2026
UR  - https://ethotechnics.org/research/the-green-dashboard
ER  -