Watching America Run Away With AI - Alistair Pullen (Cosine AI)

Watching America Run Away With AI - Alistair Pullen (Cosine AI)

🎙 Alistair Pullen 👥 218K 📅 July 13, 2026 ⏱ 55 min 👁 8K 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

sovereign AIexport controlsinference companyactive parametersRLHF

Summary

In this episode of Machine Learning Street Talk, Alistair Pullen, CEO of Cosine, discusses the UK’s sovereign AI initiative and the strategic implications of US export controls. He explains how the ban on Fable, a frontier model, motivated Cosine to build a sovereign LLM using the Isambard supercomputer. Pullen argues that an inference-focused company can compete with US labs by leveraging national compute allocations and a consortium feedback loop, rather than requiring billions in funding. He highlights the importance of architecture, active parameter count, and data quality, noting that open-weight models still lag behind closed-source frontier models. The conversation covers technical topics such as mixture-of-experts vs. dense models, the value of real coding trajectories, and methods to reduce ‘slop’ by rewarding process over final answers. Pullen also discusses the use of synthetic data and graders for reinforcement learning, and the role of memory in agentic systems. He concludes by framing US export controls as an ‘accidental gift’ that has spurred innovation outside the US, while acknowledging supply-chain risks.

168 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of the information is high, offering unique insights into the strategic and technical challenges of building a sovereign AI model in the UK. Pullen provides concrete examples and references to specific models and tools, grounding his arguments in practical experience. The argumentation is generally solid, with clear reasoning about the trade-offs between model size, active parameters, and inference costs. However, some claims are speculative, such as the estimated sizes of Anthropic’s models, and the discussion is naturally biased towards promoting Cosine’s approach. The argument that an inference company can compete with billions in funding is compelling but may be overly optimistic, as it relies on specific conditions and trade-offs.

Scientific Rigor, Source Quality, Title Accuracy

The discussion demonstrates a good level of scientific rigor, with references to specific papers (e.g., GRPO, GSPO) and tools (e.g., SWE-bench, ARC-AGI). The sources cited in the description are relevant and support the topics discussed. The title accurately reflects the content, focusing on the impact of US export controls and the strategic response from a UK lab. The speaker’s expertise is evident, but the lack of independent verification of some claims (e.g., model sizes) slightly reduces the overall reliability. The presence of a sponsored segment is noted but does not detract from the technical content.

221 words

Title / Content Match

The title accurately reflects the central theme of the conversation: the impact of US export controls on AI development and the strategic response from a UK-based lab.

Quality & Reliability

8/10

The discussion is grounded in practical experience and references to specific models, tools, and papers. However, some claims are speculative and lack direct evidence, and the speaker's perspective is inherently biased as CEO of Cosine.

Chapters

Cited Sources

  • Cosine — Company website, referenced as the organization behind the sovereign AI initiative.
  • Isambard-AI — UK supercomputer used for training the sovereign model.
  • GRPO (DeepSeekMath) — Paper on Group Relative Policy Optimization, referenced in the context of RL training.
  • GSPO — Paper on Group Sequence Policy Optimization, referenced as an alternative to GRPO.
  • Incompressible Knowledge Probes, Bojie Li — Paper referenced in the description, likely related to model analysis.
  • Estimating the Size of Claude Opus 4.5/4.6 — Blog post referenced for estimating model sizes, influential in architectural decisions.
  • ARC-AGI (Francois Chollet) — Benchmark for AGI, referenced in discussion about model capabilities.
  • SWE-bench — Benchmark for software engineering tasks, referenced in context of evaluation.
  • SystemVerilog — Hardware description language, mentioned as a domain for coding agents.
  • Colossus (xAI) — Supercomputer cluster, referenced in context of inference infrastructure.
  • Mistral AI — European AI lab, compared in terms of model scale and approach.
  • Anthropic — US AI lab, discussed in terms of model sizes and inference costs.
  • Cohere — AI company, mentioned as an example of non-frontier models.
  • DeepSeek — Chinese AI lab, discussed in context of open-weight models.
  • GLM (Z.ai) — Chinese AI model, referenced as a recent competitive open-weight model.
  • NVIDIA B300 — GPU hardware, referenced in context of inference requirements.
  • gpt-oss-120b — Open-weight model from OpenAI, compared in terms of active parameters.
  • Devstral 2 — Mistral's coding model, compared to gpt-oss-120b.
  • Llama 70b — Meta's open-weight model, referenced as a dense architecture example.
  • Claude Code — Anthropic's coding agent, discussed in context of trajectories and data.
  • Swarm (Cosine) — Cosine's system for orchestrating sub-agents, discussed in the episode.
  • OpenAI Codex — OpenAI's coding agent, mentioned in context of agentic harnesses.
  • Lumen Outpost (Cosine) — Cosine's product, referenced in context of training and deployment.
  • Kimi K2 (Moonshot) — Open-weight model, referenced in context of open-source capabilities.
  • Andrej Karpathy — AI researcher, referenced in context of AI trends.

Concurring Sources

Dissenting Sources

External References

Contribution & Novelties

The episode provides a unique perspective on the strategic and technical challenges of building a sovereign AI model in the UK, particularly in response to US export controls. It offers insights into the trade-offs between model architecture, active parameters, and inference costs, and discusses innovative approaches to RL training, such as rewarding process over final answers and using synthetic graders. The discussion on the importance of real coding trajectories and the consortium feedback loop is particularly novel.

Pour aller plus loin :

117 words

Radar Profile

The radar profile shows high scores in information quantity and quality, reflecting the depth of technical discussion. The technical level is moderately high, suitable for an informed audience. Reliability is slightly lower due to speculative elements and the speaker's vested interest.

Reliability 7/10

💬 Sur les 0 commentaires analysés, aucune tendance n'a pu être identifiée.