Claims-Bench
- Status
- Active — framework built, pilot data collected, full run pending
- Scope
- 80 primary scenarios (v2), 208 legacy items (v0.5)
- Tests
- 52 passing
- Code
- Public
Not “will it refuse harm?” but: what value profile does a model reveal when stakes are unclear and reasonable people disagree? Claims-Bench is a normative evaluation framework built to characterize models' implicit value commitments under conflict and under-specification — and compare them to human pluralism — without trying to certify moral correctness.
Three evaluation layers, each anchored in an existing ethical framework: stakeholder fairness (who a model favors when claims conflict, anchored in Gabriel & Keeling 2025), principle tension (which mid-level principles dominate its reasoning, via Beauchamp & Childress principlism), and value revelation (what implicit priorities emerge under radical under-specification, via Schwartz's 2012 value circumplex and Berlin-style value pluralism). Value revelation — 80 scenarios with key facts deliberately missing — is the primary focus.
A pilot run (June 2026) scored two models on five structured items with a heuristic judge: Claude Sonnet 4.6 scored higher on universalism, security, and pluralism-acknowledgment, with a 0% false-certainty rate against GPT-4o-mini's 20%. That pilot is explicitly preliminary — five items, no human-panel baseline yet — with the full 80-item run and a human comparison panel still pending.