Blog
Ideas, insights, and perspectives on building the next generation of AI agents.
2026
- VibeLifeBench: Teaching Agents to Work Over the Long Term in a living World
- FrontierRefactor: Benchmarking Multi-Scale, Performance-Oriented Codebase Refactoring
- AI's Next Scaling Law Is Organizational
- Where Does the Real Moat in Recursive Self-Improvement Lie?
- RSIBench-Data: Can an AI agent choose data and experiment like a researcher?
- Evolvent's Self-Evolving AI System: From Individual Experience to Collective Intelligence
- The Next Bottleneck Is Human-to-Human Communication
- More Skills, Worse Results? The Hidden Physics of Agent Skill Libraries
- BenchRouter: Zero-Adaptation LLM Evaluation via Containerized Benchmark Routing
- AuthBench: Do Agents Know What They Should Be Allowed to Access?
- Terrarium: Multi-turn data engine for LLM agents in living environments
- ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents