Changelog
What's new at Devence Lab
Articles we have published, positions we have opened, and where it started, newest first.
Why Static Interpretability Fails on Multi-Step Agentic Decision Chains
The Agent Sandbox: A Reference Architecture for Isolating Autonomous AI Systems
Runtime Monitors for Autonomous Systems: Detecting Drift and Misbehavior After Deployment
Case Study: Using Interpretability to Catch a Specific Failure Mode Before Deployment
Sparse Autoencoders: What They Reveal, and the Accuracy Tradeoffs Nobody Advertises
Attribution Graphs: A Technical Walkthrough of Circuit Tracing, and Where It Breaks
Chain-of-Thought Monitoring: A Fragile Window Into Model Cognition
devencelab.com launches
- NewThe first version of the Devence Lab website.