Agents Everywhere: OpenClaw, Codex, and the Post-Chatbot Shift
Listen to episode
About this episode
We're back! Thanks for listening! ❤️
OpenAI’s acquisition of OpenClaw signals the beginning of the end of the ChatGPT era | 2026-02-17
VentureBeat argues OpenAI’s OpenClaw move is a pivot from chatbots to agents that take actions across apps and systems. The tension is that OpenClaw’s “fast and loose” openness helped it go viral—so what happens when enterprise guardrails and safety expectations move in?
https://venturebeat.com/technology/openais-acquisition-of-openclaw-signals-the-beginning-of-the-end-of-the
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? (arXiv:2602.11988) | 2026-02-12
A new paper tests whether repo-level context files like AGENTS.md actually help coding agents finish tasks—and finds they can backfire. The punchline is a double hit: lower success rates and more than 20% higher inference cost, hinting that “more context” can mean “more confusion.”
https://arxiv.org/abs/2602.11988
AI Doesn’t Reduce Work—It Intensifies It | 2026-02-09
Simon Willison highlights research suggesting AI can increase work intensity instead of easing it, especially when productivity gains mask burnout. The provocative angle is managerial: if AI boosts throughput, how do organizations prevent that extra capacity from turning into an always-on expectation?
https://simonwillison.net/2026/Feb/9/ai-intensifies-work/
SWE-rebench Leaderboard | n.d.
SWE-rebench is trying to solve a messy problem in agent evaluation: benchmarks go stale, and models get “contaminated” by training on the tasks. The leaderboard format makes it feel like a live sport—but the real question is whether continuously refreshed tasks can keep results honest as models ship faster.
https://swe-rebench.com/
Tidbits/extras:
Introducing Claude Opus 4.6 | 2026-02-05
Anthropic says Claude Opus 4.6 upgrades its top-tier model for longer, more reliable agentic coding and better performance in large codebases. The headline-grabber is a 1M-token context window in beta—prompting the question of whether bigger memory finally means fewer brittle, lost-in-the-middle failures.
https://www.anthropic.com/news/claude-opus-4-6
Introducing GPT-5.3-Codex-Spark | 2026-02-12
OpenAI is pitching GPT-5.3-Codex-Spark as a speed-and-feedback upgrade that makes coding agents feel
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity