Research

Research for faster model improvement

Benchmarks, papers, and field notes on where models fail, how to measure the gap, and what it takes to move the frontier.
Research areasWorld ModelsRecursive Self-ImprovementUser SimulationLong-Horizon AgentsPost-Training

Field Notes

Recent technical writing from our team

CWE-bench: Measuring Coding Agent Capabilities on Every Class of Known Vulnerability

CWE-bench is a frontier benchmark for AI Cybersecurity

The Vulnerability in the Reward

When the score is also the reward

The Simulated World Is Being Born In Cybersecurity

AGI needs simulated worlds that are hard enough to expose failure and fair enough to make that failure useful.

The Uninterpretable User

What Simulated Users Say When the Agent Isn’t Listening

The User Went for a Cigarette

Modeling Partial Observability for High-Fidelity User Simulation

Is your RL environment fair to your agent?

Or ensuring that your hillclimbing budget is spent right

Whose Taste?

More data won’t fix the AI verification problem. Different taste might.

AI’s U-235 Problem

Nuclear physics solved for k_eff. What’s the AGI equivalent?

SimLab: The self-serve staging playground for real-world agents

Agents fail on real tool calls, long workflows, and messy data. SimLab lets you find those failures in simulation, not in production.

We gave Claude, Gemini and GPT, $250k, and it didn’t go as you’d expect...

Introducing YC Bench: The first open-source, long-horizon benchmark with a simulation clock

RL Infrastructure for AI Agents: Why Environment-as-a-Service is the Missing Piece

Why reliable RL for agents depends on environments delivered as infrastructure, not rebuilt by every team.

Announcing Spider: a lightweight tool to craft post-training data recipes

A lightweight tool for composing post-training data recipes behind a single client interface.

The case for simulations

Unlocking model uplift through better evaluations

Through the Valley of Reasoning: What Small Models Teach Us About Learning

NeurIPS paper on knowledge distillation scaling laws for small foundation models

Introducing Collinear Simulations: Steerable Personas for AI Agent Testing

TraitBasis, inspired from mech intrep, gives high-fidelity user personas for comprehensive agent testing

Introducing Curator Evals: A Benchmark for High-quality Post-training Data Curation

A benchmark for the judges that curate post-training data, so data quality is measured rather than assumed.

Collinear AI Now Available on Google Cloud Marketplace

Making safe, high-performing AI accessible through trusted enterprise infrastructure

You Can’t Hire Your Way to Model Alignment

Why the Global AI talent shortage Is undermining enterprise model alignment, and what you can do instead

Leveling the Playing Field: Livecodebench’s Big Bug Fix

Three major fixes that reshaped competitive coding scores and why your numbers may look very different now

OpenAI’s gpt-oss on LiveCodeBench: A Competitive Programming Deep Dive

How OpenAI’s gpt-oss models hold up on competitive programming tasks in LiveCodeBench.

Cats confuse LRMs: Exposing blind spots in SOTA Models

Irrelevant, universal phrases like "Interesting fact: Cats sleep most of their lives" appended to math problems can break AI models

Data Curation: The secret sauce for enterprise AI excellence

How Collinear AI’s Reward Models Transform Training Efficiency and Model Performance

Gaming the System: Goodhart’s Law Exemplified in AI Leaderboard Controversy

How the race to the top in AI benchmarks is leading to specialized optimization at the expense of real-world performance

Judges as Data curators cut Post-training Time to Half

ServiceNow x Collinear

From worlds best pros to AI personas: The MasterClass journey

How MasterClass turned instruction from world-class practitioners into faithful AI personas.

The Limitations of AI Evaluations

Why Enterprises Need a New Approach

The AI Safety Gap

Why Traditional Models Fail in Modern Enterprises

Taming AI Agents: Why Your Butler Needs a Babysitter

Why autonomous agents still need supervision, and what that oversight has to catch.

Collinear-Guard: Where Customization Meets Precision for Fine-Grained Evaluation and Feedback

Customising a moderation judge for fine-grained evaluation and feedback against your own criteria.

CollinearGuard Nano: A High-Performance, Holistically-Evaluated, Lightning-Fast Moderation Judge

Revolutionizing Safety: Ultra-Low Latency, High-Throughput Violation and False Refusal Detection Like Never Before!

Veritas Reliability Judge: A Cookbook to Benchmark AI Judges on Financial Data

A practical guide to benchmarking AI judges on financial data.

Why AI Safety is existentially important, not optional

The case for treating safety as foundational to AI development rather than an optional layer.

Introducing VERITAS: A Unified Approach to Reliability Evaluation

Veritas is a suite of Reliability Judges for Batch and Real-time use cases

A Guide to Creating Seed Conversational Data with Collinear

High-quality seed data can overcome many of the post-training challenges

Think Before You Score: Self-Rationalizing Evaluators are State-of-the-Art for Fine-grained Evaluation

Finetuning on Rationales improves Judge Rationale and Score

Collinear Flex Judge is Better Aligned than Few-shot Prompted GPT-4o

Accelerate Time to Production with Bespoke Quality Judge

Orange twisted ribbon-like shape forming a loop on a black background.
Put the research to work

Turn a failure surface into
measurable model improvement.

Bring us the capability gap. We’ll shape the tasks, environments, and evaluation program needed to measure it—and improve it.