MLOps.community
MLOps.community

Sandboxing, Agent Harnesses, and Agent Teamwork

19 June 2026 1:19:53 Demetrios

Listen to episode

About this episode

Shahram Anver is the co-founder and CEO of Cleric, the company building the first self-learning AI SRE: an autonomous agent that investigates production incidents, performs root cause analysis, and builds a continuously updated knowledge graph of your infrastructure. Before Cleric, Shahram led MLOps at Gojek, where he built the CaraML platform serving billions of monthly predictions.


In this conversation we go from harness to human: how much harness is too much, why the investigation agent is now the easiest part of an AI SRE stack, and what 2027 looks like when your coding agent and your SRE agent actually have to work together at 2AM.


What we cover:


- Defining the agent harness. Everything between the user and the model. Why Claude Code and Codex are themselves harnesses, and where the next layer of abstraction lives.

- How much harness is too much. Why rigid custom Kubernetes tools made sense in 2023 and were a dumb idea by 2024. The trade-off between safety and reaching better performance.

- "Nondeterminism in a strong box." Shahram's mental model for sandboxing agents once you let them run bash and Python.

- Why the investigation agent is the easiest part now. Testing, environments, verification, and learning are the durable hard problems. Search is cheap. Knowing the agent is right is hard.

- The subtle-incident problem. Sev-1s are easy because they are binary. The latency spike at 20 percent is what breaks every generic agent.

- Documenting failure modes. Running every trace through another LLM, labeling mistakes, and deciding when to fix it in the harness vs let the agent course-correct itself.

- The trust paradox. Why one engineer told Shahram he prefers when the agent is wrong. What happens to your spidey sense once the agent is right most of the time.

- Manager vs craftsman. Why the future operator is both. Reviewing decision traces instead of reams of generated code.

- The hopper. Shahram's term for the work queue you load before bed and wake up to almost-done. The new dopamine loop of vibe-coding a PR into open source.

- The 1-on-1 problem and the stand-up problem. Setting policies with a single agent, and the much harder problem of agents working together across teams.

- SRE agent vs coding agent. Coding is narrow input with infinite tasks. SRE is horizontal breadth-first search with a finite set of problem shapes.

- Will Cleric just become a skill? How the founders decide what is durable vs what to never spend an iota of time on. Why "do not fine-tune" was rule number one.

- 2027. Your coding agent ships a bad config at 2AM. Your SRE agent reverses it, watches the metrics, and writes you a report you read with coffee.


If you build, run, or are on-call for production systems, or if you are trying to figure out how to give agents enough rope to be useful without enough rope to take down prod,

Want to find AI jobs?

Join thousands of AI professionals finding their next opportunity

We respect your inbox. Unsubscribe at any time.

© 2026 MLOps.community. All rights reserved.

Common Questions

Frequently asked questions

Quick answers about how DevFound's AI matching, resumes, and referrals work.

DevFound's AI Copilot ingests your profile, goals, and live job data to deliver curated matches in seconds. Every match includes a resume variant, suggested referrals, and interview prep so you can act immediately. The more feedback you provide, the sharper the Copilot becomes.

AI-led job searches shrink the hours spent sifting through boards and formatting resumes. DevFound pairs automation with your personal outreach, so you reserve energy for interviews and negotiation. Traditional networking still matters, but AI gives you a lift before you even send a message.

Modern AI roles expect comfort with production-grade code, data fluency, and practical ML tooling. The strongest candidates pair deep technical chops with storytelling—translating model impact to product, GTM, and exec partners. Continuous learning keeps you ahead as stacks evolve.

DevFound rewards active seekers. Keep your profile fresh, respond to match quality prompts, and enable alerts so you never miss a role. The AI prioritizes companies and teams that align with your feedback, accelerating both introductions and interview invites.

High-density tech hubs continue to host the deepest AI talent pools, yet distributed teams are catching up fast. Use DevFound filters to hone in on onsite, hybrid, or fully remote roles and watch openings expand across time zones.

DevFound aggregates thousands of remote AI openings and flags the nuances—core hours, async culture, and visa needs—up front. The Copilot also recommends how to position your distributed work experience so hiring managers know you can thrive on a remote team.