Models & Research
Researchers propose judging models on selective trust, not resistance alone — A new preprint introduces MIST, a 1,000-item benchmark that renders each reasoning item under four matched conditions, clean, misleading, correct-context, and irrelevant-context, and SCOPE, a preference-training method that balances all four inside a standard DPO objective.