What we’re
learning.
Research, ideas, and updates from Kaon AI and Kaon Labs. One shared collection of everything we’re learning and building.

Online Evaluation at Scale: How We Evaluate 100+ Model Variants per Week
How Session AB measures sustained interactions with real users, helps us evaluate over a hundred model variants a week, and relates to engagement in user-level A/B tests.
Read the article
Reading a roleplay model before it speaks: what differs after post-training
After roleplay post-training, where do Post-trained and Baseline differ before generation? A J-space diagnostic across 32,768 online conversation states, with selected examples and clear limits on what they show.
Read the article
Online Autoresearch: Running Autoresearch with Real Online Feedback
We ran an agent through 7 SFT experiments end-to-end — from data filtering to online ELO evaluation — using real user feedback as the only reward signal. The best run lifted online win rate from 56.0% to 60.1%.
Read the article
Let the world grow because of you
Generative systems break the old user-item frame. We want software that carries a person's influence forward, so the worlds it creates can grow around the people inside them.
Read the article
What personalization means to us
Why personalization in generative systems is not memory, retrieval, or tone control, but the ability to build and update a useful model of a person.
Read the article
Offline Evaluation Is Dead: How Consumer AI Learns from Millions of Real User Decisions
Once consumer AI products reach scale, offline evaluation isn't just "not good enough"—it's irrelevant to production decisions. This conclusion comes from running systems at tens of millions of users, processing billions of conversations and behavioral signals every month. We've seen this pattern repeatedly in real user environments: models that top public creative writing benchmarks often underperform in actual usage.
Read the articleElsewhere.
Kaon in the news.
