Research home

Papers

Browse the full research library.

Ten papers per page, with summaries, diagrams, formulas, concepts and citations kept close to the source.

All papers

21 research papers / Page 1 of 3

GPT-6 Sol: Reasoning, Coding Agents, and the Economics of Frontier AI visual
LLMs / Agentic AI/13 min read/Sep 22, 2026

GPT-6 Sol

A source-grounded deep dive into GPT-6 Sol covering model behavior, API/runtime design, benchmarks, effort economics, safety, limitations, and production agent architecture.

Read this to understand what GPT-6 Sol actually exposes, what OpenAI has disclosed about its development, how its benchmark results should be interpreted, and how to design a production agent around its reasoning/cost trade-offs.

OpenAI

Read paper
Grok 4.7: The Shift to Long-Horizon Agent Work visual
LLMs / Agentic AI/12 min read/Sep 21, 2026

Grok 4.7 — Long-Horizon Agents

A source-grounded technical walkthrough of Grok 4.7 training changes, runtime controls, multimodal input, tool use, context management and long-horizon agent engineering.

Read this to understand what SpaceXAI says changed in Grok 4.7, what is public versus proprietary, and how its API/runtime design maps to long-running agents.

SpaceXAI

Read paper
Grok 4.7 Benchmarks: Vendor Tables, Independent Runs, and Harness Effects visual
AI Evaluation / Benchmarks/14 min read/Sep 21, 2026

Grok 4.7 — Benchmarks

A provenance-first benchmark dossier for Grok 4.7, preserving the official comparison matrix and separating it from Artificial Analysis measurements.

Read this when you need to interpret Grok 4.7 launch numbers without mixing vendor runs, independent runs, benchmark versions, effort tiers or harnesses.

SpaceXAI, Artificial Analysis

Read paper
Grok 4.7 in Production: Agent Architecture, Cost, Tools, and Safety visual
ML Systems / Agentic AI/13 min read/Sep 21, 2026

Grok 4.7 — Production

A systems-oriented deployment guide covering the Grok 4.7 API, agent loop, caching, long-context economics, rate limits, observability and safety boundaries.

Read this to design a production-grade Grok 4.7 agent service rather than treating the model as a stateless chat endpoint.

SpaceXAI

Read paper
Claude Opus 5.5 in Production: The Real Story Is Cost per Task visual
AI Models / Production/11 min read/Sep 24, 2026

Opus 5.5 — Production Economics

Production economics view: compare unit price, task trajectory, code volume, review effort, and verification cost.

Read this when deciding whether Opus 5.5 changes the economics of a production agent or coding workflow.

Anthropic, Sonar, Endor Labs

Read paper
Claude Opus 5.5: What It Does Well, Where It Still Falls Short visual
AI Models / Review/12 min read/Sep 24, 2026

Opus 5.5 — Strengths & Limits

Strengths and limitations with explicit separation between Anthropic claims, independent measurements, and model-review observations.

Read this for the uncomfortable details that a vendor benchmark page cannot answer by itself.

Anthropic, Sonar, Endor Labs

Read paper
Claude Opus 5.5: What Changed and Why This Model Feels Different visual
AI Models/9 min read/Sep 22, 2026

Opus 5.5 — What Changed

A plain-English model overview focused on what changed from Opus 5 and what those changes mean in daily use.

Read this for the clean product-level picture before diving into benchmarks or deployment details.

Anthropic

Read paper
Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers visual
AI Agents/11 min read/Sep 10, 2026

AgentActionBench

A process-level benchmark for reliable research-agent execution.

Useful for designing evals for coding and research agents where execution, reproducibility, and traceability matter.

Hanhua Hong, Yizhi Li, Luu Gia Huy

Read paper
Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification visual
AI Agents/12 min read/Sep 10, 2026

Magenta

Verification-guided mathematical reasoning with Lean 4.

Useful for verifier-guided agents, formal correctness boundaries, and test-time compute trade-offs.

Joshua Ong Jun Leang, Haonan Li, Zheng Zhao

Read paper
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI visual
AI Agents/14 min read/Sep 17, 2026

ScientistTwo

Autonomous ML research with explicit experiment, ablation and review loops.

Understand what a production autonomous research workflow needs beyond a single agent.

Jaehyun Nam, Jinsung Yoon, Yanzhou Pan

Read paper