Sandboxing, Agent Harnesses, and Agent Teamwork
Listen to episode
About this episode
Shahram Anver is the co-founder and CEO of Cleric, the company building the first self-learning AI SRE: an autonomous agent that investigates production incidents, performs root cause analysis, and builds a continuously updated knowledge graph of your infrastructure. Before Cleric, Shahram led MLOps at Gojek, where he built the CaraML platform serving billions of monthly predictions.
In this conversation we go from harness to human: how much harness is too much, why the investigation agent is now the easiest part of an AI SRE stack, and what 2027 looks like when your coding agent and your SRE agent actually have to work together at 2AM.
What we cover:
- Defining the agent harness. Everything between the user and the model. Why Claude Code and Codex are themselves harnesses, and where the next layer of abstraction lives.
- How much harness is too much. Why rigid custom Kubernetes tools made sense in 2023 and were a dumb idea by 2024. The trade-off between safety and reaching better performance.
- "Nondeterminism in a strong box." Shahram's mental model for sandboxing agents once you let them run bash and Python.
- Why the investigation agent is the easiest part now. Testing, environments, verification, and learning are the durable hard problems. Search is cheap. Knowing the agent is right is hard.
- The subtle-incident problem. Sev-1s are easy because they are binary. The latency spike at 20 percent is what breaks every generic agent.
- Documenting failure modes. Running every trace through another LLM, labeling mistakes, and deciding when to fix it in the harness vs let the agent course-correct itself.
- The trust paradox. Why one engineer told Shahram he prefers when the agent is wrong. What happens to your spidey sense once the agent is right most of the time.
- Manager vs craftsman. Why the future operator is both. Reviewing decision traces instead of reams of generated code.
- The hopper. Shahram's term for the work queue you load before bed and wake up to almost-done. The new dopamine loop of vibe-coding a PR into open source.
- The 1-on-1 problem and the stand-up problem. Setting policies with a single agent, and the much harder problem of agents working together across teams.
- SRE agent vs coding agent. Coding is narrow input with infinite tasks. SRE is horizontal breadth-first search with a finite set of problem shapes.
- Will Cleric just become a skill? How the founders decide what is durable vs what to never spend an iota of time on. Why "do not fine-tune" was rule number one.
- 2027. Your coding agent ships a bad config at 2AM. Your SRE agent reverses it, watches the metrics, and writes you a report you read with coffee.
If you build, run, or are on-call for production systems, or if you are trying to figure out how to give agents enough rope to be useful without enough rope to take down prod, this one is for you.
Links and Resources:
- Cleric: https://cleric.ai/
- Shahram Anver on LinkedIn: https://www.linkedin.com/in/shahramanver/
- Willem Pienaar (Cleric CTO, creator of Feast): https://www.linkedin.com/in/willempienaar/
- Cleric launches the first self-learning AI SRE: https://cleric.ai/blog/cleric-launches-the-first-self-learning-ai-sre
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity