Operational certificates for world models

Beyond Multimodal
Alignment:

Certifying Physical Language through Response Substitution
and Ordered Execution

1New York University  ·  2Carnegie Mellon University  ·  3Columbia University

4.5×closer than a wrong-surface bridge
19/19held-out surfaces separate
14/16registered checks pass

The capability ladder

A representation is not one score.
It is a profile.

Each rung changes exactly one element of physical-language use while the response vocabulary stays fixed. A certificate that returns a single verdict cannot tell them apart.

  1. Attribute access

    Can one observation recover a physical variable?

    prior work
  2. Response substitution

    Do two sensors denote the same response function?

    certified · 4.5× · 19/19
  3. Fusion closure

    Does the joint predictive law actually improve?

    beats both; diagonal ties
  4. Ordered execution

    Does the phrase preserve the intervention law?

    certified at 25× budget

Demonstration I  ·  Cluster Haptic

One word, two independent sources

A robot scans 118 surfaces. Audio and acceleration are compiled by separate models on disjoint query panels — no matching loss, no entity identifier — then decoded by one frozen chart. Pick a held-out surface: the two sensors land on the same response. Pick a wrong surface for either one and they part.

Decoded response · surface-normal channel
audio · P₀ accel. · P₁ unsealed truth
0.5161same-surface bridge
1.1145population substitution
2.3104wrong-surface bridge
r = 0.63bridge distance vs. true error

Demonstration II  ·  live simulation in your browser

One phrase, two orders

A unit mass on a spring, a dashpot and a plastic slider. The slider remembers: once a pulse yields it, the rest position moves. So A then B does not settle where B then A does. Run them and watch. This is the exact system, integrated live at Δt = 0.005.

plastic slip p0.000
displacement x0.000
force f0.000
stateelastic
Scored readout the last 60 steps, after the drive ends
r(AB) r(BA) run both orders to measure Δ

The certificate

An instrument, not a verdict

The gate has a prerequisite: before any learned inference is scored, the exact chart coordinate must execute a held-out program better than the population response. At the pre-registered budget it does not — so the gate refuses, correctly. Train the executor to convergence and the same chart clears it. Swap the executor's factorization and the gate stops it again.

When the gate is cleared

Oracle NMSE against executor updates. The registered 1.2k budget misses; the crossing sits between 6k and 12k.

A property of the executor

Absolute NMSE vs. an entity-blind predictor, log scale. The program-level executor is better on words it was fitted on, and collapses on an unseen adjacency.

Fusion closes, diagonal ties

Held-out joint energy score. Fusion improves on both modalities; its diagonal restriction does slightly better, the sole remaining gate failure.

The paper

Beyond Multimodal Alignment

Certifying Physical Language through Response Substitution and Ordered Execution.

BibTeX
@article{tan2027certifying,
  title  = {Beyond Multimodal Alignment:
            Certifying Physical Language through Response Substitution and Ordered Execution},
  author = {Tan, Kaizhen and Xu, Xin and Tao, Siru and Li, Yixiao
            and Hong, Hanzhe and Feng, Yang and Du, Heqing},
  year   = {2027}
}