Briefing

Intercom’s Scaling Architecture: Vitess, Elasticsearch, and Multi‑Tenant Guardrails

hosting
by Ryan Sherlock ·

Configure AWS SQS fair queues for each tenant and enable Vitess sharding to isolate customer DB traffic.

What to do now

Enable Vitess sharding for new tenants and set up SQS fair queues to isolate spikes.

Summary

Intercom’s platform now handles over 150,000 customer requests per second and 70,000 asynchronous requests per second at peak, with daily peaks of more than five million conversations and 100 million comments. The company migrated its source‑of‑truth database to Vitess on PlanetScale, creating 128 shards that process roughly two million DB requests per second and support ten million cache reads per second. Elasticsearch clusters have grown to 650 TB of storage, 1.7 trillion documents, and support over 40,000 requests per second, with online index reshaping enabled by partitioning, dual‑writing, and feature flags. Multi‑tenant isolation is enforced through AWS SQS fair queues, overflow queues, and application‑level guardrails, ensuring a single customer’s spike does not degrade platform latency for others. The AI Agent Fin adds new scaling challenges, routing across multiple model providers and protecting customer workloads from lower‑priority tasks. Together, these changes give Intercom practical scaling levers and reduce the need for disruptive architectural rewrites.

Key changes

  • Intercom now processes 150k customer requests per second and 70k async requests per second at peak
  • Database layer uses Vitess on PlanetScale with 128 shards, handling ~2M DB requests per second and 10M cache reads per second
  • Elasticsearch clusters store 650 TB, 1.7 trillion docs, and support >40k requests per second
  • Multi‑tenant isolation achieved via AWS SQS fair queues, overflow queues, and application guardrails
  • AI Agent Fin requires routing across multiple model providers and protecting customer workloads from lower‑priority tasks
  • Online index reshaping in Elasticsearch uses partitioning by customer ID, dual‑writing, backfilling, and feature flags
  • Vitess provides native sharding, query routing, online schema changes, connection pooling, and resharding primitives
  • The architecture allows independent scaling of database capacity per customer, including dedicating a shard to a single customer

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting