Briefing

Cognition Labs Raises $10B Series C, Projects $1B ARR

ai-dev
DeepSeek

Patch your inference stack to use Perplexity Unigram tokenizer and DeepSeek V4‑Pro attention for lower cost, and integrate LangChain Deep Agents v0.6 to shrink checkpoint size.

What to do now

Implement Perplexity Unigram tokenizer, switch to DeepSeek V4‑Pro for long‑context inference, and migrate agent pipelines to LangChain Deep Agents v0.6 to reduce storage.

Summary

Cognition Labs, the fastest‑growing independent agent laboratory, closed a $10 billion Series C round in September, propelling its valuation to a 2.5‑fold increase over the past eight months. The funding boost positions the company as the largest independent lab in the agent‑AI space and underpins its projection of over $1 billion in annual recurring revenue by year‑end. The round was led by a consortium of venture firms that see Cognition’s portfolio—now featuring enterprise names such as Exa and Modal—as a key driver of next‑generation AI services. AINews has migrated to the lab’s Latent Space platform, and Cognition plans to unpack the financial and technical milestones in an upcoming podcast.

On the technology front, Cognition’s EAGLE 3.1 release introduces speculative decoding robustness and a longer‑context acceptance length, enhancing model reliability for complex tasks. Perplexity’s open‑source Unigram tokenizer slashes CPU usage by 5–6× and delivers a 63‑microsecond latency on 514 tokens, while Qwen 3.5 achieves 580 tokens per second for agentic workloads on TokenSpeed. MaxSim v2 adds backpropagation, boosting inference speed by 10.33× on H200 GPUs and 11.94× on A100s. DeepSeek V4‑Pro’s hybrid attention mechanism shrinks a 1‑million‑token key‑value cache to roughly 10 % of the previous version and cuts single‑token FLOPs to 27 %, dramatically reducing memory and compute footprints.

LangChain’s Deep Agents v0.6 further tightens resource usage by introducing Delta Channels, which compress a 200‑turn coding session checkpoint from 5.3 GB to 129 MB. The update also adds computer‑use capabilities in Fleet and a Context Hub for versioned skills. LangSmith Engine automates the evaluate‑diagnose‑fix loop, while Xiaomi MiMo’s stochastic weight averaging and hierarchical cache cut caching costs by about 80 %. Collectively, these releases lower inference costs and improve agent memory fit, enabling more scalable deployment of autonomous systems.

The convergence of robust funding, cutting‑edge model optimisations, and memory‑efficiency tools signals a maturation of the agent‑AI ecosystem. With a projected $1 billion ARR and a suite of performance‑enhancing releases, Cognition Labs is poised to lead the industry toward more cost‑effective, high‑throughput autonomous applications, while its partners and clients anticipate faster, cheaper, and more reliable AI services.

Key changes

  • EAGLE 3.1 improves speculative decoding robustness, stabilizes hidden‑state feedback, reduces attention drift, focuses on long‑context acceptance length.
  • Perplexity Unigram tokenizer cuts CPU utilization 5–6×, 63 µs at 514 tokens, zero heap allocations.
  • Qwen3.5 TokenSpeed hits 580 tokens/s for agentic workloads via joint optimization.
  • DeepSeek V4‑Pro uses hybrid attention, compressed sparse attention, 1 M‑token KV cache 10 % of V3.2, single‑token inference FLOPs 27 %.
  • LangChain Deep Agents v0.6 introduces Delta Channels, reduces checkpoint storage from 5.3 GB to 129 MB, adds computer use in Fleet, Context Hub, LangSmith Engine.
  • Trajectory platform enables post‑deployment learning, funded $15 M, supports 397 B‑parameter model deployment with FP8/NVFP4 quantization on autoscaled H100 infra.
  • MaxSim v2 backprop speeds 10.33× on H200, 11.94× on A100 versus naïve PyTorch.
  • Gemini Embedding 2 white paper introduces native multimodal embeddings over text, image, audio, and video.

Affects

internal

Source angles · 2 perspectives

AI News
Independent angle

not much happened today

Open
Latent Space
Independent angle

Cognition Series C, Inference Optimizations, and Agent Tool Releases

Open

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting