Back to media
Human In the Loop · EP 20

Nobody Knows Who Built Ox Alpha

Podcast EpisodeAugust 25, 2026
PodcastAI ModelsAI AgentsAI Privacy
In this episode

Ep 20: Nobody Knows Who Built Ox Alpha

Ox Alpha appeared on OpenRouter on August 20, 2026. The anonymous model is free during its preview, accepts text, images, and video, supports tool calls, and has a 1,048,576-token context window.

No company has claimed it. Developers have theories. None of them is confirmed.

The privacy term is less mysterious. OpenRouter says the anonymous provider retains prompts and completions. That makes Ox Alpha useful for public and synthetic tests, not private code or customer data.

Signal or Noise

  1. Ox Alpha, SIGNAL with a hard caveat. The blind preview gives builders a real model to test. The hidden provider, missing technical report, small community benchmark, and retained prompts rule out trust by default.
  2. NVIDIA AVO, SIGNAL. NVIDIA's agent system completed all 183 levels in the ARC-AGI-3 public set using Claude Opus 5. NVIDIA warns that the roughly 30 percent model baseline and 100 percent system result are not a controlled comparison.
  3. OpenAI Astra pause, SIGNAL. OpenAI paused frontier reinforcement-learning work for two weeks after preliminary evidence that Astra may meet its Critical cybersecurity threshold. Its largest planned frontier RL run remains on hold.
  4. Reconstruction benchmark, SIGNAL. Seven frontier models reached only about 3 to 15 percent on a blind research idea test. A multi-agent method reached about 23 to 42 percent. The paper is a preprint and uses an LLM judge.
  5. DeepSeek Harness rc.8, SIGNAL for builders. DeepSeek can now install Codex and Claude Code as subagent components. The release remains a developer preview and changes its storage format incompatibly.

Ship It or Skip It

  1. Agent System CI: Turn agent failures into repeatable tests of the model, memory, tools, permissions, retries, time, and cost. Ship a focused workflow or skip another general evaluation platform?
  2. R&D Idea Tournament: Let several agents propose and challenge research ideas before a human expert decides what deserves a trial. Ship the workshop or skip the autonomous-research story?

Closing takes

Oscar: Anonymous testing removes brand bias. Retained prompts remove any excuse for sending private code.

Matt: The NVIDIA result says model rankings are incomplete. Completed work depends on the full system.

Your hosts

  • Oscar Gallo: AI Engineer and entrepreneur. I live in the intersection of engineering and businesses.
  • Matt Wozniak: Serial Builder and relentless executor. I come from the lens of what works and what doesn't.

Listen now

Sources

Your move

Like how we think about AI?

Human In the Loop is me and Matt thinking out loud. Putting that thinking to work inside a real company is the day job. If you're a founder or team trying to ship AI that survives production, let's talk.

Free 30-minute call · No pitch, just a plan · No commitment until you say go

Keep going

More from the archive