Metrics for Your AI Workers: Build a Personal Agent Benchmark

When you run a team, you don’t decide who to keep based on their resume from two years ago. You watch what they actually ship on your work, and you promote the ones who deliver and let go of the ones who don’t. I want the same thing for the models I use as agents. And public leaderboards don’t give it to me. Why public benchmarks don’t tell you what you need A leaderboard tells you how a model did on someone else’s tasks, scored on someone else’s rubric, often on problems that leaked into training data....

June 30, 2026 · 5 min · joor0x

The Mythos/Fable Blockade: How US AI Exceptionalism Builds Its Own Rivals

This is an opinion piece. My read on the news, not reporting. Disagree freely. Last Friday (June 12, 2026), the Trump administration forced Anthropic to pull Mythos 5 and Fable 5 — three days after launch. The official reason: a “national security” risk over a minor jailbreak. The real reason looks simpler: if we can’t control who uses it, nobody gets it. I think that’s a mistake. Here’s why....

June 17, 2026 · 2 min · joor0x

Selecting an Open-Source DB for Financial Time Series

Choosing Your Data Engine: More Than Just Code When your algorithms depend on processing high-frequency data streams, or when you’re building ML models that need fast access to vast historical context, the time series database isn’t just a component – it’s the bedrock of your operation. A bottleneck here means missed opportunities, flawed analysis, or outright system failure. I’ve spent time evaluating the options because getting this wrong has consequences, especially when real capital or critical infrastructure is on the line....

April 6, 2025 · 8 min · Josep Oriol Carné