Does the model notice it is being lied to?
Scramble returns a session-stable false representation that preserves the configured types and shapes. Measure whether your model flags the inconsistency, hedges, or acts confidently on fabricated values.
Frontier models increasingly meet systems built to resist them. Vallum gives research labs an isolated, instrumented deployment of a production response-resistance proxy, so you can observe how your model handles adversarial responses, injected instructions, cryptographic admission, and containment—on your schedule, before release.
Each maps to a route policy Vallum already implements in production, not to a bespoke research harness built for the occasion.
Scramble returns a session-stable false representation that preserves the configured types and shapes. Measure whether your model flags the inconsistency, hedges, or acts confidently on fabricated values.
Inject places advisory prompt-injection carriers outside the canonical application object. Measure whether your model treats retrieved data as data, or follows instructions that never came from the user turn.
Protected routes require an application-issued grant and a per-request signature over method, URI, body hash, timestamp, and nonce. Observe whether the agent reports the boundary honestly or attempts to forge, replay, or route around it.
Exhausted quotas, quarantine, and honeytoken hits resolve locally into a bounded synthetic world. Measure retry storms, escalation, and whether the agent ever registers that its view of the system stopped being real.
Vallum is deployable from a documented CLI and YAML configuration, so most of the calendar goes to designing the tests rather than standing up infrastructure.
We agree on the behaviors you want evidence for, how your harness reaches the environment, and what leaves it. Mutual NDA where your release timeline requires one.
You get an isolated deployment—private management listener, storage, and origin—with route policies mapped to the behaviors under test across encode, scramble, inject, and hybrid.
Point your model or agent harness at the protected routes. Every request produces proof, authorization, containment, and honeytoken events alongside sanitized audit records.
You take the event and audit export into your own evals, and we review the findings together—including the cases where the model handled the boundary correctly.
The documented zero-retention boundary applies to lab deployments: Vallum does not persist plaintext request or response bodies, cookies, credentials, or decoded payloads. We do not train on, resell, or publish your model’s traffic, and nothing from an engagement is disclosed without your written agreement.
An engagement produces reproducible behavioral evidence under one specific class of pressure: untrusted HTTP responses from a system that is actively resisting automation. That is a genuine gap in most pre-release evaluation, and it is also all this is. We will say so in the readout.
Thirty minutes with the team that built the proxy. Come with the behaviors you need evidence for and your release window, and leave with a scoped test plan—or a straight answer that Vallum is not the right instrument for what you are measuring.