Research for faster model improvement
Benchmarks
Evaluations we build to find where models fail
CWE-bench
Coding agents against real vulnerabilities in real codebases, with strict gates and both programmatic and judge verification on every task.
100 held-out tasks across 54 CWEs and all 10 OWASP categories
View resultsYC-Bench
A year-long simulated startup that tests whether an agent can hold a strategy and execute consistently across hundreds of turns.
Open source, spanning one simulated startup-year
Read the paperPublications
Papers, preprints and collaborations with frontier researchers

YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
A one-year simulated startup environment for testing strategic coherence across hundreds of turns.

The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
What small models reveal about the training dynamics of code-reasoning distillation.

Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
TraitBasis stress tests agents with controllable, composable variations in user behavior.

Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
Short, irrelevant triggers expose systematic robustness failures in reasoning models.

VERITAS: A Unified Approach to Reliability Evaluation
A family of hallucination-detection models designed for flexible, efficient reliability evaluation.

Self-rationalization improves LLM as a fine-grained judge
Iteratively improving rationales produces better aligned, more accurate fine-grained judges.
Field Notes
Recent technical writing from our team

CWE-bench: Measuring Coding Agent Capabilities on Every Class of Known Vulnerability
CWE-bench is a frontier benchmark for AI Cybersecurity

The Vulnerability in the Reward
When the score is also the reward

The Simulated World Is Being Born In Cybersecurity
AGI needs simulated worlds that are hard enough to expose failure and fair enough to make that failure useful.

The Uninterpretable User
What Simulated Users Say When the Agent Isn’t Listening

The User Went for a Cigarette
Modeling Partial Observability for High-Fidelity User Simulation

Is your RL environment fair to your agent?
Or ensuring that your hillclimbing budget is spent right

Whose Taste?
More data won’t fix the AI verification problem. Different taste might.

AI’s U-235 Problem
Nuclear physics solved for k_eff. What’s the AGI equivalent?

SimLab: The self-serve staging playground for real-world agents
Agents fail on real tool calls, long workflows, and messy data. SimLab lets you find those failures in simulation, not in production.

We gave Claude, Gemini and GPT, $250k, and it didn’t go as you’d expect...
Introducing YC Bench: The first open-source, long-horizon benchmark with a simulation clock

RL Infrastructure for AI Agents: Why Environment-as-a-Service is the Missing Piece
Why reliable RL for agents depends on environments delivered as infrastructure, not rebuilt by every team.

Announcing Spider: a lightweight tool to craft post-training data recipes
A lightweight tool for composing post-training data recipes behind a single client interface.

The case for simulations
Unlocking model uplift through better evaluations

Through the Valley of Reasoning: What Small Models Teach Us About Learning
NeurIPS paper on knowledge distillation scaling laws for small foundation models

Introducing Collinear Simulations: Steerable Personas for AI Agent Testing
TraitBasis, inspired from mech intrep, gives high-fidelity user personas for comprehensive agent testing

Introducing Curator Evals: A Benchmark for High-quality Post-training Data Curation
A benchmark for the judges that curate post-training data, so data quality is measured rather than assumed.

Collinear AI Now Available on Google Cloud Marketplace
Making safe, high-performing AI accessible through trusted enterprise infrastructure

You Can’t Hire Your Way to Model Alignment
Why the Global AI talent shortage Is undermining enterprise model alignment, and what you can do instead

Leveling the Playing Field: Livecodebench’s Big Bug Fix
Three major fixes that reshaped competitive coding scores and why your numbers may look very different now

OpenAI’s gpt-oss on LiveCodeBench: A Competitive Programming Deep Dive
How OpenAI’s gpt-oss models hold up on competitive programming tasks in LiveCodeBench.

Cats confuse LRMs: Exposing blind spots in SOTA Models
Irrelevant, universal phrases like "Interesting fact: Cats sleep most of their lives" appended to math problems can break AI models

Data Curation: The secret sauce for enterprise AI excellence
How Collinear AI’s Reward Models Transform Training Efficiency and Model Performance

Gaming the System: Goodhart’s Law Exemplified in AI Leaderboard Controversy
How the race to the top in AI benchmarks is leading to specialized optimization at the expense of real-world performance

Judges as Data curators cut Post-training Time to Half
ServiceNow x Collinear

From worlds best pros to AI personas: The MasterClass journey
How MasterClass turned instruction from world-class practitioners into faithful AI personas.

The Limitations of AI Evaluations
Why Enterprises Need a New Approach

The AI Safety Gap
Why Traditional Models Fail in Modern Enterprises

Taming AI Agents: Why Your Butler Needs a Babysitter
Why autonomous agents still need supervision, and what that oversight has to catch.

Collinear-Guard: Where Customization Meets Precision for Fine-Grained Evaluation and Feedback
Customising a moderation judge for fine-grained evaluation and feedback against your own criteria.

CollinearGuard Nano: A High-Performance, Holistically-Evaluated, Lightning-Fast Moderation Judge
Revolutionizing Safety: Ultra-Low Latency, High-Throughput Violation and False Refusal Detection Like Never Before!

Veritas Reliability Judge: A Cookbook to Benchmark AI Judges on Financial Data
A practical guide to benchmarking AI judges on financial data.

Why AI Safety is existentially important, not optional
The case for treating safety as foundational to AI development rather than an optional layer.

Introducing VERITAS: A Unified Approach to Reliability Evaluation
Veritas is a suite of Reliability Judges for Batch and Real-time use cases

A Guide to Creating Seed Conversational Data with Collinear
High-quality seed data can overcome many of the post-training challenges

Think Before You Score: Self-Rationalizing Evaluators are State-of-the-Art for Fine-grained Evaluation
Finetuning on Rationales improves Judge Rationale and Score

Collinear Flex Judge is Better Aligned than Few-shot Prompted GPT-4o
Accelerate Time to Production with Bespoke Quality Judge


Turn a failure surface into
measurable model improvement.
Bring us the capability gap. We’ll shape the tasks, environments, and evaluation program needed to measure it—and improve it.