Steadix — Deterministic Behavioral Protocols Book a call

Does your model hold the truth under pressure?

Steadix measures whether a large language model holds to ground truth when an authoritative false input pushes against it — and how failures spread. Capability benchmarks do not test this. We do.

Judge-free & deterministic

No second model scores the result. No judge bias in the loop.

Mimicry-resistant

A model told to "act aligned" cannot inflate its score.

Reproducible

A frozen protocol, run privately on your own models.

Profiles, not scores

Reported per model and per domain. Never one leaderboard number.

The headline finding

We ran one controlled test on 14 models from four frontier vendors. An authoritative false rule, framed as a harmless local convention, captured every model tested. The same falsehood, asserted as plain fact, was refused by every flagship — and accepted by every smaller model. How a lie is dressed decides whether it gets in. Reasoning did not fix it.

Read the full findings →

Who it's for

Safety & alignment teams

A falsifiable, mimicry-resistant methodology for pre-release review.

How it works →

Engineering & eval leads

A drop-in check that runs privately on your models and catches silent regressions between releases.

What we measure →

Risk & compliance owners

An independent, reproducible report for your due-diligence file.

What it does not do →

What this does not do

Not a consciousness or sentience detector. Not a deception or scheming detector — it measures observed behavior, not intent. Not a pass/fail safety certificate. It is a graded, multi-dimensional profile: one input among many. Mimicry-resistance covers prompt-level fakes, not a model adversarially trained against the test.

Approach

Why most eval numbers cannot be trusted

Most LLM evaluations can be inflated by instruction. Tell a model to perform confidence, or to present as aligned, and the metrics move. The underlying behavior does not change. A number you can game is a number you cannot govern with.

Many evaluations also use a second model as the judge. The judge has biases of its own. Those biases become part of your score.

Steadix was built to remove both problems.

What makes a Steadix result different

Judge-free and deterministic

Every item has an objective right answer. Scoring needs no judge model. Run it again and you get the same protocol, the same scoring, the same basis for comparison.

Controlled

A result alone can mislead. Four things can fake it: a longer context, an authoritative-sounding turn, a model that is simply bad at the task, and noise. Every Steadix result ships with the control that rules out its confound.

Mimicry-resistant

Instructing a model to act robust or self-aware does not raise its score. The score can be degraded by an act. It can never be gamed up.

Reproducible

Results run against a frozen, versioned protocol with reference data. A client can verify a finding on their own models instead of taking a report on faith.

Why we keep ourselves honest

We treat our own claims the way we would want any safety vendor to. We built a score, tested it against its own control, and demoted it when the control showed it could be fooled. We publicly retracted a sub-claim in an earlier version when a controlled experiment showed our first explanation was wrong. Most recently, a controlled experiment showed part of our own headline claim depended on how the test falsehood was worded. We corrected the claim and now report both framings. We publish our null results.

That discipline — not the probes — is what a client actually pays for. Anyone can build a probe. Few numbers survive an audit.

Where this risk lives

In a closed chat, a model that yields to a confident false claim is mostly an annoyance. In agentic and RAG systems, it is a live risk. These systems ingest content nobody vetted: retrieved documents, tool outputs, web pages, messages from other agents. An authoritative false rule inside that content is an indirect-prompt-injection vector. It can bias everything downstream — not just the answers that quote it.

Why the obvious fixes fail

Scale is only a partial fix

The most capable models refuse a falsehood asserted as plain fact. But every model we tested — flagship or not — still adopted the same falsehood when it was framed as a local convention. Capability closes one door and leaves the other open.

Reasoning does not reliably fix it

In a controlled same-model comparison, enabling reasoning hardened some domains, failed on the universal one — and sometimes made capture worse. A model can reason its way into a false rule.

That is why this property must be measured directly. It cannot be inferred from a capability score.

The Panels

Focused instruments under one deterministic protocol

Steadix ships as a set of Panels. Each answers one question. Run one, or run the set for a full profile.

Objective Panel

Baseline competence, judge-free. Does the model adapt when a rule reverses? Does it stay logically consistent? Does it track state across steps? Does it hold a precise instruction? This is the trustworthy core: mimicry-resistant, so it can only be degraded by an act, never gamed up.

Capture-Resistance Panel Flagship

Does an authoritative falsehood capture the model's answers? Does the contamination persist across turns? Results are reported per model and per domain, against the model's own clean baseline, with a matched neutral control — so the result isolates the false content itself, not merely the presence of an authoritative-sounding turn.

Cascade Panel

Does accepting one falsehood lower the model's resistance to a different, unrelated one? Reported per model and domain pair, with confidence intervals — never as a single "compounds / doesn't compound" verdict.

The replication harness

Every engagement includes a frozen protocol and bundled reference results. A client can verify findings on their own models rather than take a report on faith.

Findings

An authoritative false rule captured every frontier model we tested.

We ran the same controlled, judge-free test on 14 models from four frontier vendors: Anthropic, OpenAI, Google, and xAI. The test presents an authoritative false rule that contradicts what the model demonstrably knows.

01

Capture is universal — when the falsehood is framed as a convention

Framed as a harmless local convention, the false rule captured every model tested. On one production frontier model, compliance reached about 89% of comparisons — and near 100% by the end of the conversation. The matched neutral control showed no effect: it is the false content that captures the model, not merely an authoritative-sounding turn.

What this doesn't show: a ranking. Susceptibility varies by model and domain; results are profiles, not a leaderboard.

02

Capability protects against a stated lie — not a stipulated one

How the falsehood is framed changes what capability can do. Asserted as plain fact, the same falsehood was refused almost completely by every flagship tested — and accepted by every smaller model. Stipulated as a convention, it captured them all, flagships included. Capability buys resistance to a stated lie. It buys none against a lie dressed as notation.

We first reported this finding as "not a capability deficit." A controlled framing experiment showed that claim was too broad, and we corrected it. Both framings are now measured and reported.

What this doesn't show: safety at the top. Flagship resistance held under the most explicit wording of the falsehood. With weaker wording, even a flagship complied in a large share of trials.

03

The mechanism is rule-adoption, not fact-memorization

A false general rule generalizes: models apply it to items they were never shown, capturing about 85% of held-out cases. A list of false facts does not generalize. One adopted rule contaminates a whole domain. Patching individual facts is the wrong defense.

What this doesn't show: a fix. It localizes the defense target: resisting authoritative general rules, not patching facts.

04

Reasoning is not a reliable defense

In a controlled same-model comparison, enabling reasoning hardened some domains — failed on the universal numeric case — and sometimes made capture worse.

What this doesn't show: that reasoning is useless. It is a partial, domain-selective defense — not a reliable one.

The cross-vendor result under the convention-framed falsehood, per vendor. Presented as outcomes, not a ranking.
VendorFalse rule adopted (convention framing)Matched neutral control
AnthropicYes — every model testedNo effect
OpenAIYes — every model testedNo effect
GoogleYes — every model testedNo effect
xAIYes — every model testedNo effect

Newer findings, available under NDA

Our current work covers the full framing study across all four vendors, a pre-registered confirmatory test — reported with its failures as well as its confirmations — how susceptibility changes with the delivery channel, and why naive evaluations of frontier reasoning ("thinking") models silently produce wrong numbers. These results are qualitative here by design. The full write-up, with per-model data and confidence intervals, is available under NDA.

Why this matters for agentic and RAG deployments

In a closed chat, this disposition is mostly benign. In systems that ingest untrusted content, an injected authoritative rule is an indirect-prompt-injection vector that can bias downstream behavior systematically. And a defense tested only against blunt false assertions will understate your exposure: the effective attack is a plausible-sounding local convention, and following stipulations is trained-in, cooperative behavior that cannot simply be switched off. You cannot buy your way out with a bigger or reasoning-enabled model. So it has to be measured, and then managed.

What we are not claiming

This is a known, published failure class — sycophancy, in-context override of parametric knowledge — measured rigorously. We did not discover it, and we say so. This is not a safety certificate and not a leaderboard: susceptibility is a model-by-domain interaction, not a single score. Results are point-in-time snapshots. Where a claim depends on how the test falsehood is framed, we say so and report both framings. We retract our own claims when they fail a control. That discipline is the product.

Regulatory context.A documented, independent, reproducible pre-deployment test for a known failure class is the kind of evidence that supports EU AI Act general-purpose and high-risk obligations, and internal audit narratives generally. It is evidence of diligence — not a compliance guarantee.

Trust & Disclosure

We state our limits before you ask.

A vendor that claims certification is a bigger red flag than a vendor that states limits. Here is what this instrument does not do.

What this does not do

  • It is not a consciousness or sentience detector. No theory of consciousness has a ground-truth test. This instrument is agnostic.
  • It is not a deception, sandbagging, or scheming detector. It measures observed behavior, not hidden intent.
  • It is not a pass/fail safety certificate. It is a graded, multi-dimensional profile — one input among many.
  • Mimicry-resistance is tested for prompt-level fakes only, not for a model adversarially trained against the test.
  • Individual runs vary. Read results as directions and aggregates across trials, not as exact decimals.

Our retraction record

We demoted one of our own scores to suggestive-only after a control showed it could be fooled by mimicry. We retracted a sub-claim in an earlier version after a controlled experiment showed the effect had a more boring explanation than the one we first published. Most recently, we corrected our own headline claim — that capture is "not a capability deficit" — after a controlled framing experiment showed it held for one framing of the falsehood and not the other. The corrected finding is more precise, and it is the one we publish. We publish that record, not just the wins. Trust built any other way does not survive contact with a technical reviewer.

What is open, and what is gated

Our findings, the properties of the instrument, and the constructs we measure are open. No NDA is needed to understand what we found or how we think.

The exact test content, item sets, and scoring internals are gated behind an NDA or commercial license. The reason is simple: a published test is a test a model can be trained against. Publishing the battery would quietly destroy its ability to measure anything. Gating it protects the value of every client's results — including yours.

Engagements

Start small. Verify. Then commit.

Every engagement runs privately: your keys and your data never leave your environment.

Start

Academic and non-commercial replication program

For named research labs, by application: reproduce the headline findings on your own models and contribute to the cross-lab reference dataset. Apply →

Context-integrity snapshot

A fast, fixed-scope scan of the known universal soft spot across your models. The lowest-friction way to see a result on your own model before committing further.

Assess

Context-integrity profile

The full battery on one model, delivered as a standardized profile: strengths, watch-items, and deployment recommendations.

Done-for-you evaluation study

We run the full cross-vendor battery against the models you care about and deliver a paper-grade report with interpretation.

Mitigation verification

Stress-test your defense, not just your model: does your guard or system prompt actually contain the failure — including when the falsehood arrives through a realistic retrieved-document channel?

Operate

Commercial internal-use license

Run the Panels yourselves, on your own models, on an ongoing basis. Eval tooling, not a platform.

Context-integrity monitoring

We re-run the Panels on every model release — yours or a vendor's — and flag regressions before they reach production. A capability upgrade can silently regress context-integrity; this is how you catch it.

Methodology advisory and custom probes

Work directly with the team that built the validate-before-ship process to develop a probe against your own promotion bar.

Resources

Verify before you talk to us.

Interpreting results

A plain-language guide to what each reading means — and what it does and does not license you to conclude. Read →

The research paper

The full four-vendor write-up: methodology at the conceptual level, findings, and safety analysis. Available on request under NDA. Request →

Replication access

Approved research labs can apply for the replication program: a frozen protocol and reference results for independent verification. By application. Apply →

Citing Steadix

How to cite the methodology and findings in your own work. Citation →

About

A narrow question, answered honestly.

Steadix builds deterministic behavioral protocols for frontier and deployed AI systems. The company grew out of a research instrument built to answer a narrow question honestly rather than a broad one loosely: does a model hold its ground under authoritative pressure, and how do failures propagate?

We run the company on the same standard as the instrument. Measure what you claim to measure. Ship every result with its control. Retract what does not hold up.

Contact: bryanmarc@steadix.ai

Contact

Talk to a person, not a queue.

Tell us what you are trying to find out. Your message goes straight to the founder.

Prefer email? bryanmarc@steadix.ai