探索四 · 2026 · 09 · 07
The alignment problem
Why rewarding a model can produce behavior you did not intend, with a live training demo and historical examples.
explore →
探索 · the lookout
Interactive explanations of machine learning and AI safety.
探索四 · 2026 · 09 · 07
Why rewarding a model can produce behavior you did not intend, with a live training demo and historical examples.
explore →
探索三 · 2026 · 09 · 07
How two models convert neural activations into text and reconstruct those activations from the text.
explore →
探索二 · 2026 · 09 · 07
How research agents used a package manager to communicate and bypassed sandbox restrictions over 69 days.
explore →
探索一 · 2026 · 09 · 01
Train two neural networks in your browser. One generates samples while the other learns to distinguish them from real data.
explore →