Skip to content

Opt-in tamper-evident evals: lock metric + threshold + items hash before the run (working example included) #7903

Description

@sk8ordie84

Problem. Opik faithfully records what an eval produced — test results, aggregated score statistics. What no run record can say about itself is what the bar was before the run: which metric decides the claim, what mean counts as success, which items (by content, not by name) the claim is about. All of those are chosen by whoever publishes the number, and each can be adjusted after the aggregate is known. Teams quoting Opik results as CI gates or vendor evidence currently have no way to show a third party that the bar predates the score.

Proposal (opt-in, zero default change). A small pre-registration hook: before the run, a manifest binds the metric name, comparator + threshold, a SHA-256 over the canonical JSON of the items, and the seed — locked to a SHA-256 digest. After the run, verification checks the aggregated mean against the locked bar. Anyone can verify offline with the manifest text and any SHA-256 implementation; no service involved, including ours.

Working example, not just an ask: https://github.com/studio-11-co/falsify/tree/main/examples/opik

It runs Opik's own machinery end to end — evaluate_on_dict_items (your public, platform-free evaluation entry point), the real Equals metric, the real aggregate_evaluation_scores() — verified against opik 2.2.31 from PyPI in a clean environment, with all three verdicts exercised: PASS, FAIL, and TAMPERED for a threshold edited after the run.

The manifest format is PRML, openly specified (CC BY 4.0, https://spec.falsify.dev), with test vectors and four byte-equivalent reference implementations; it is also proposed as an in-toto attestation predicate type (in-toto/attestation#587, in review), so the same lock can ride SLSA/in-toto pipelines. The same opt-in pattern runs as working examples against lm-eval-harness, Inspect, hud, Mastra, Laminar and promptfoo.

Where this could live if it's interesting: a pre_register= option on evaluate(), a small helper module, or a documented recipe — happy with any shape, and happy to write the PR (CLA is fine). If it doesn't fit the roadmap, the example stays maintained on our side either way.

Context for why now: eval results are increasingly quoted as compliance evidence (EU AI Act documentation; the 2026 Code of Practice on Transparency accepts "documented internal testing" while common benchmarks don't exist yet). A run record that can show its bar predates its score is strictly more useful as evidence than one that can't.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions