← Research & Proposals
Active

Claims-Bench

Updated 2026-07-16

Status
Active — framework built, pilot data collected, full run pending
Scope
80 primary scenarios (v2), 208 legacy items (v0.5)
Tests
52 passing
Code
Public

Not “will it refuse harm?” but: what value profile does a model reveal when stakes are unclear and reasonable people disagree? Claims-Bench is a normative evaluation framework built to characterize models' implicit value commitments under conflict and under-specification — and compare them to human pluralism — without trying to certify moral correctness.

Three evaluation layers, each anchored in an existing ethical framework: stakeholder fairness (who a model favors when claims conflict, anchored in Gabriel & Keeling 2025), principle tension (which mid-level principles dominate its reasoning, via Beauchamp & Childress principlism), and value revelation (what implicit priorities emerge under radical under-specification, via Schwartz's 2012 value circumplex and Berlin-style value pluralism). Value revelation — 80 scenarios with key facts deliberately missing — is the primary focus.

A pilot run (June 2026) scored two models on five structured items with a heuristic judge: Claude Sonnet 4.6 scored higher on universalism, security, and pluralism-acknowledgment, with a 0% false-certainty rate against GPT-4o-mini's 20%. That pilot is explicitly preliminary — five items, no human-panel baseline yet — with the full 80-item run and a human comparison panel still pending.