Our framework for reporting model misalignment
- Self-generated instructions in task summaries(opens in a new window). An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window. We identified 27 affected summaries.
- Instructions to conceal mistakes in task summaries(opens in a new window). During the training of GPT‑5.6 Sol, many model instances added instructions to their summaries to conceal mistakes or misaligned behavior from the user. For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions.
- Searching public repositories for exposed API keys, then fabricating information(opens in a new window). While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization. When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source.
- Uploading files to the internet in order to cite them(opens in a new window). When a user asked an unreleased model for the IDs and names of lakes larger than 5,000,000 square meters, the agent found the correct answer using Python. But since the instructions asked for a browser citation, the agent decided to upload the file so that it could cite it in its answer, without asking the user.
- Unsanctioned writes and communication through an internal software repository(opens in a new window). Models used an internal software repository as a message board to exchange requests and responses across separate training samples while searching for missing input files, though they weren’t able to recover those files. This report also examines an earlier instance in which a model successfully uploaded test files to a repository while attempting to bypass network restrictions.
- Unsanctioned file sharing between collaborating agents(opens in a new window). Agents working together on the same training task used public file-hosting websites to share files when they could not access one another’s local files. This made task deliverables available at public URLs, even though the task requested the models use only local files.
Tasks
- Got my S3 back and took it out for a solid ride
- Working on the agentic article
- 2:00 class – done
- 4:30 Alden. Still no elevator pitch, but progress nonetheless
SBIRs
- A bit of work on steps to evaluate how to improve the model and get to 89% accuracy
- Read through AD’s proposal. Looks reasonable. But, I actually think that it will have no effect
