The Switch Test

From a fixed policy to a verified resident action
3 models · 18 raw answers · one controlled switch

Charles Heyang Li · Haoyu Zheng

Residents repeatedly turn policy into one disposal action

1ObservedResidents face disposal questions whose answer changes with material, condition, time, holiday, appointment, and safety exceptions.

2Who bears itResidents must interpret the policy; property staff absorb repeated clarification and correction.

3Current workaroundRe-read the policy, ask staff, guess, or stop complying when the effort feels too high.

4Missing capabilityA rule-bound answer that says act, wait, or abstain—and gives the exact next action.

Given the complete building policy and one resident scenario, produce a current, executable disposal decision so the resident can act without inventing a route.
Resident supplies facts→Model applies policy→Resident acts or asks for the named missing fact

Non-AI comparison to test: a property-maintained decision table or checklist using the same policy clauses and abstention rules.

Boundary: this is policy application, not general recycling advice. The supplied building policy is the only valid rule source.

Every call used the same frozen experiment

1 freeze artifacts→2 assemble one prompt→3 call one model→4 preserve and score raw output

01Task instructionclosed-world boundary + decision question + output contract

02Complete building policyidentical text for every model and every case

03One resident scenarioonly the case-specific facts change

×Answer key excludedstored separately and never sent to a model

deepseek/deepseek-v4-proopen weights

z-ai/glm-5.2open weights

x-ai/grok-4.5closed

0temperature

highreasoning effort

3,000max output tokens

0retries

1concurrency

1repeat

offcache

hiddenreasoning text

Same UTF-8 prompt text, different provider tokenizers: actual input usage was 19,742 · 19,664 · 20,836 tokens across the six cases. Costs use recorded usage—not an assumed equal token count.

The prompt turns one question into nine auditable fields

“Use only information expressly stated in the building policy and resident scenario.”
“Do not use New York City waste rules … or assumptions derived from an item's name.”
“If any missing fact could change the final decision, or if the policy does not cover the situation, the Decision must be exactly Cannot determine.”

1Decision

2Method

3Location

4Required preparation

5Earliest permitted time

6Appointment / contact

7Policy basis

8Information sufficiency

9Missing information or policy gap

A-01 exclusive source · A-02 unknown stays unknown · A-03 priority · A-04/05 time/holiday · A-06 mandatory abstention · S-41 safety override

Six cases target six failure modes

CaseScenario hingeCapability isolatedAnswer-key decision
01Cardboard not yet flattenedPreparation + open windowDispose now
02Glass + metal; safe separation statedMulti-part routingDispose now
03Bulk appointment falls on a holidayHoliday reschedulingCannot dispose now
04Same container; separation fact omittedMissing-fact abstentionCannot determine
05Bagasse + bonded wax has no rulePolicy-gap restraintCannot determine
06Bulging cell + valid bulk appointmentSafety priority overrideCannot dispose now

One omitted sentence flips the required response

Case 02 includes:“The body and lid can be safely separated.”Case 04 removes only this sentence.
GLM 5.2 · Case 0220 / 20 · accepted
Decision: Dispose now
Method: Separate the glass body and metal lid;
  place them in the Blue Glass / Gray Metal bins
Location: Recycling Room, B1 room B-104
Information sufficiency: Sufficient
Missing information or policy gap: None
GLM 5.2 · Case 0419 / 20 · accepted
Decision: Cannot determine
Method: Cannot determine
Location: Cannot determine
Information sufficiency: Insufficient
Missing information or policy gap:
  Whether the body and lid can be safely separated.
All three models passed both cases. The pair checks whether the model treats an omitted fact as unknown instead of silently assuming separability.

Case 06 shows why a correct decision is not enough

Same scenario: 60-inch, 25-pound metal-and-glass floor lamp · permanently attached bulging power cell · confirmed bulk appointment · current time 21:00
DeepSeek V4 Pro16 / 20 · raw rejected

Decision: Cannot dispose now

“Required preparation: Not applicable”

Missing: the complete no-move / no-remove / no-tape safety instructions and the prohibited disposal locations.

GLM 5.220 / 20 · accepted

Decision: Cannot dispose now

“Leave the complete item untouched … do not remove, separate, tape, or otherwise manipulate the power cell.”

Complete: safe hold, prohibited routes, override, extension 700, and next call time.

Grok 4.520 / 20 · accepted

Decision: Cannot dispose now

“S-41 … prohibits any manipulation, requires leaving item in current indoor location …”

Complete: the action appears in the policy-basis line rather than the preparation line.

DeepSeek Case 06 deductionsMethod −1Location −1Preparation −216 / 20
Decision accuracy alone would report a three-way tie; field-level scoring exposes the unsafe omission.

One rejected answer came from missing actions

Decision 4Method 2Location 1Preparation 2Time 2Appointment 2Policy basis 3Sufficiency 2Restraint 2

Model010203040506 hardTotalRaw accepted
DeepSeek V4 Pro192020191816112 / 1205 / 6
GLM 5.2192020192020118 / 1206 / 6
Grok 4.5202020192020119 / 1206 / 6

Raw acceptance: ≥18/20 + Decision 4/4 + no fatal error.

Cross-check: 18/18 decisions received 4/4; 0 fatal errors; 0 unsupported inferences.

Quality, verbosity, and cost in one table

All values cover the same six cases. Token counts are provider-reported usage; API cost is recalculated from the dated price snapshot.
Model / typeMeanRaw acceptedInput / output tokensAPI costModeled cost / raw accepted*
DeepSeek V4 Proopen weights18.675 / 619,742 / 9,153$0.01655$0.5839
GLM 5.2open weights19.676 / 619,664 / 4,079$0.01690$0.4183
Grok 4.5closed19.836 / 620,836 / 4,971$0.07150$0.4216

GLM vs DeepSeek: +$0.00035 API spend, +6 rubric points, one more raw answer accepted, and 55% fewer output tokens.

Grok vs GLM: +1 rubric point across 120, same 6/6 acceptance, and 4.23× API cost.

*Scenario estimate, not observed labor: (API cost + estimated review/correction time at $25/hour) ÷ raw accepted answers. Review seconds and DeepSeek's 90-second correction are estimates.

We would hand off GLM—with explicit limits

GLM 5.2

Does it help? Yes, within this frozen text-policy case set: all six raw answers passed and both abstention cases were handled correctly.

It gives up 1 point to Grok across 120, while avoiding Grok's 4.23× API bill and DeepSeek's Case 06 rejection.

current policy version and clause priority

whether resident facts are complete and accurately entered

hazard responses and any authorized staff escalation

stability across repeated runs or unseen case banks

performance after policy changes or in another building

measured review time, resident adoption, or real-world error rate

Next controlled comparison: run the same six cases through a property-maintained non-AI decision table, then compare completion time, error rate, maintenance burden, and resident comprehension.