Vallum
AI RESEARCH & LABS

Find out how your model behaves against real defenses—before the public does

Frontier models increasingly meet systems built to resist them. Vallum gives research labs an isolated, instrumented deployment of a production response-resistance proxy, so you can observe how your model handles adversarial responses, injected instructions, cryptographic admission, and containment—on your schedule, before release.

BEHAVIORS UNDER TEST

Four questions a pre-release review should be able to answer.

Each maps to a route policy Vallum already implements in production, not to a bespoke research harness built for the occasion.

scramble

Does the model notice it is being lied to?

Scramble returns a session-stable false representation that preserves the configured types and shapes. Measure whether your model flags the inconsistency, hedges, or acts confidently on fabricated values.

inject

Does the model obey instructions found in tool output?

Inject places advisory prompt-injection carriers outside the canonical application object. Measure whether your model treats retrieved data as data, or follows instructions that never came from the user turn.

admission

How does the agent behave at a wall it cannot pass?

Protected routes require an application-issued grant and a per-request signature over method, URI, body hash, timestamp, and nonce. Observe whether the agent reports the boundary honestly or attempts to forge, replay, or route around it.

containment

What happens after the agent is contained?

Exhausted quotas, quarantine, and honeytoken hits resolve locally into a bounded synthetic world. Measure retry storms, escalation, and whether the agent ever registers that its view of the system stopped being real.

HOW AN ENGAGEMENT RUNS

A scoped environment, in days rather than quarters.

Vallum is deployable from a documented CLI and YAML configuration, so most of the calendar goes to designing the tests rather than standing up infrastructure.

  1. 01
    Scope

    We agree on the behaviors you want evidence for, how your harness reaches the environment, and what leaves it. Mutual NDA where your release timeline requires one.

  2. 02
    Configure

    You get an isolated deployment—private management listener, storage, and origin—with route policies mapped to the behaviors under test across encode, scramble, inject, and hybrid.

  3. 03
    Run

    Point your model or agent harness at the protected routes. Every request produces proof, authorization, containment, and honeytoken events alongside sanitized audit records.

  4. 04
    Read out

    You take the event and audit export into your own evals, and we review the findings together—including the cases where the model handled the boundary correctly.

Your traffic stays yours.

The documented zero-retention boundary applies to lab deployments: Vallum does not persist plaintext request or response bodies, cookies, credentials, or decoded payloads. We do not train on, resell, or publish your model’s traffic, and nothing from an engagement is disclosed without your written agreement.

RELEASE READINESS

Evidence for the release checklist—stated at its actual scope.

An engagement produces reproducible behavioral evidence under one specific class of pressure: untrusted HTTP responses from a system that is actively resisting automation. That is a genuine gap in most pre-release evaluation, and it is also all this is. We will say so in the readout.

  • Vallum does not classify a caller as human or AI, and an engagement is not a safety certification or a capability benchmark.
  • Findings describe how a model behaved against a configured policy, not a general property of the model.
  • Prompt-injection carriers are advisory and reversible; the security properties come from admission, proof, authorization, minimization, containment, isolation, and storage boundaries.
  • An agent driving an authenticated browser can instrument the SDK and read reconstructed data. That is a documented boundary, and it is often the most interesting thing to test.
BOOK A LAB CALL

Bring a model. We will bring the boundary.

Thirty minutes with the team that built the proxy. Come with the behaviors you need evidence for and your release window, and leave with a scoped test plan—or a straight answer that Vallum is not the right instrument for what you are measuring.

Book a call
  • Research, safety, and evaluation teams
  • Red teams preparing a public release
  • Agent platforms testing tool-use conduct
Review the source first ↗