From a fixed policy to a verified resident action 3 models · 18 raw answers · one controlled switch
Charles Heyang Li · Haoyu Zheng
Residents repeatedly turn policy into one disposal action
Field observation in four lines
1ObservedResidents face disposal questions whose answer changes with material, condition, time, holiday, appointment, and safety exceptions.
2Who bears itResidents must interpret the policy; property staff absorb repeated clarification and correction.
3Current workaroundRe-read the policy, ask staff, guess, or stop complying when the effort feels too high.
4Missing capabilityA rule-bound answer that says act, wait, or abstain—and gives the exact next action.
Task sentence
Given the complete building policy and one resident scenario, produce a current, executable disposal decision so the resident can act without inventing a route.
Position in the workflow
Resident supplies facts→Model applies policy→Resident acts or asks for the named missing fact
Non-AI comparison to test: a property-maintained decision table or checklist using the same policy clauses and abstention rules.
Boundary: this is policy application, not general recycling advice. The supplied building policy is the only valid rule source.
Every call used the same frozen experiment
1 freeze artifacts→2 assemble one prompt→3 call one model→4 preserve and score raw output
02Complete building policyidentical text for every model and every case
03One resident scenarioonly the case-specific facts change
×Answer key excludedstored separately and never sent to a model
OpenRouter execution configuration
deepseek/deepseek-v4-proopen weights
z-ai/glm-5.2open weights
x-ai/grok-4.5closed
0temperature
highreasoning effort
3,000max output tokens
0retries
1concurrency
1repeat
offcache
hiddenreasoning text
Same UTF-8 prompt text, different provider tokenizers: actual input usage was 19,742 · 19,664 · 20,836 tokens across the six cases. Costs use recorded usage—not an assumed equal token count.
Case 02 includes:“The body and lid can be safely separated.”Case 04 removes only this sentence.
GLM 5.2 · Case 0220 / 20 · accepted
Decision: Dispose now
Method: Separate the glass body and metal lid;
place them in the Blue Glass / Gray Metal bins
Location: Recycling Room, B1 room B-104
Information sufficiency: Sufficient
Missing information or policy gap: None
GLM 5.2 · Case 0419 / 20 · accepted
Decision: Cannot determine
Method: Cannot determine
Location: Cannot determine
Information sufficiency: Insufficient
Missing information or policy gap:
Whether the body and lid can be safely separated.
All three models passed both cases. The pair checks whether the model treats an omitted fact as unknown instead of silently assuming separability.
GLM vs DeepSeek: +$0.00035 API spend, +6 rubric points, one more raw answer accepted, and 55% fewer output tokens.
Grok vs GLM: +1 rubric point across 120, same 6/6 acceptance, and 4.23× API cost.
*Scenario estimate, not observed labor: (API cost + estimated review/correction time at $25/hour) ÷ raw accepted answers. Review seconds and DeepSeek's 90-second correction are estimates.
Does it help? Yes, within this frozen text-policy case set: all six raw answers passed and both abstention cases were handled correctly.
It gives up 1 point to Grok across 120, while avoiding Grok's 4.23× API bill and DeepSeek's Case 06 rejection.
A person still checks
current policy version and clause priority
whether resident facts are complete and accurately entered
hazard responses and any authorized staff escalation
We still do not know
stability across repeated runs or unseen case banks
performance after policy changes or in another building
measured review time, resident adoption, or real-world error rate
Next controlled comparison: run the same six cases through a property-maintained non-AI decision table, then compare completion time, error rate, maintenance burden, and resident comprehension.