Max Agency Podcast: Why the best agents are simpler than you think
Patch your agent design to keep full context.
Patch your agent design to keep full context.
Summary
Sierra, the conversational AI platform behind many Fortune 20 companies, reveals its agent strategy in a June 25, 2026 podcast. The company runs multiple models in parallel—Claude, Gemini, and GPT‑class models—trusting each where it excels, such as a model that transcribes thick UK accents but hallucinate during silence. Sierra adopts outcome‑based pricing for high‑value work like closing a sale, while using seat‑based pricing for commodity tasks such as balance checks. The default architecture is one agent per brand, keeping the full customer history, context, and capabilities in a single agent to avoid blind triage and task agents. A no‑code layer compiles user prompts into agent code, and the platform includes a PCI‑certified stack for voice payments. Sierra also emphasizes that the model’s perceived “dumbness” is often a mismatch between the task and the model’s strengths.
The podcast discusses how Sierra’s modular voice architecture separates thinking, listening, and talking in parallel, and how the company builds a memory‑first primitive that is not a breakout memory company. It also touches on the importance of outcome‑based pricing, the pitfalls of multi‑agent systems that mirror an org chart, and the future of “more AI” as the solution to all problems.
Key changes
- Sierra runs multiple models in parallel (Claude, Gemini, GPT) for different tasks
- Uses outcome‑based pricing for high‑value work
- Defaults to one agent per brand to avoid blind triage
- Provides a no‑code layer that compiles prompts into agent code
- Includes a PCI‑certified stack for voice payments
- Separates thinking, listening, and talking in parallel
- Implements a memory‑first primitive that is not a breakout memory company
- Highlights pitfalls of multi‑agent systems mirroring org charts