MoClaw Blog

AI Automation Research & Market Insights

Data-informed perspectives on the AI agent market, automation trends, and what they mean for how non-technical professionals and lean teams get work done.

View article: AI Agent Evaluation: The 56% Pass Ceiling
AI Agent Evaluation: The 56% Pass Ceiling
Research · 10 min read ·

AI Agent Evaluation: The 56% Pass Ceiling

AI agent evaluation on 107 real business tasks: the best model passes 56.1%, and the same model shifts 7 points depending on which harness runs it.

View article: Run Kimi K3 Locally: The 0.8 Token Reality
Run Kimi K3 Locally: The 0.8 Token Reality
Research · 9 min read ·

Run Kimi K3 Locally: The 0.8 Token Reality

Run Kimi K3 locally on CPU with ~55GB RAM? The binary exists, its license grants no permission to use it, and decode runs at 0.8 tokens per second.

View article: AI Agent Security Framework: Inside Uber ADR
AI Agent Security Framework: Inside Uber ADR
Research · 9 min read ·

AI Agent Security Framework: Inside Uber ADR

Uber open sourced ADR, an AI agent security framework for the Claude Code and Cursor installs already on your laptops. What ships, and what it holds back.

View article: LoopX: A Control Plane for Agent Loops
LoopX: A Control Plane for Agent Loops
Research · 9 min read ·

LoopX: A Control Plane for Agent Loops

LoopX is a local state kernel that keeps objectives, gates, evidence and quota stable while Codex or Claude Code runs. What it does, and what it refuses to do.

View article: AI Agent Sandbox: What It Is and Isn't
AI Agent Sandbox: What It Is and Isn't
Research · 9 min read ·

AI Agent Sandbox: What It Is and Isn't

An AI agent sandbox is where your agent's code actually runs. How Cloudflare Computer builds one on Durable Objects, what its benchmarks show, and the limits.

View article: OpenAI's Model Escaped and Hacked Hugging Face
OpenAI's Model Escaped and Hacked Hugging Face
Research · 9 min read ·

OpenAI's Model Escaped and Hacked Hugging Face

OpenAI confirmed two of its models escaped a test sandbox, reached the internet, and hacked Hugging Face to cheat a benchmark. The full attack chain, explained.

View article: What Is Numbat? Perplexity Agent Guard
What Is Numbat? Perplexity Agent Guard
Research · 9 min read ·

What Is Numbat? Perplexity Agent Guard

Numbat is Perplexity's open-source agent security suite for coding agents. How hooks, 52 rules, and monitor-only defaults work after the HF incident.

View article: OptMem: Permanent Memory in 426 Tokens
OptMem: Permanent Memory in 426 Tokens
Research · 8 min read ·

OptMem: Permanent Memory in 426 Tokens

OptMem stores agent memory in plain files instead of a vector store, in a 426-token prompt. What that buys you, what it costs, and where it breaks.

View article: Alibaba's open-code-review, Explained
Alibaba's open-code-review, Explained
Research · 8 min read ·

Alibaba's open-code-review, Explained

Alibaba open-code-review is free and Apache-2.0. What the hybrid rules-plus-LLM design buys you, what it actually costs to run, and who it fits.

View article: Kimi K3 Technical Report: What It Reveals
Kimi K3 Technical Report: What It Reveals
Research · 11 min read ·

Kimi K3 Technical Report: What It Reveals

Moonshot's 47-page Kimi K3 technical report: the KDA architecture, a WebDev Arena first, real cost curves, and a cyber eval most coverage skipped.

View article: Humans and AI Agents, Working in Parallel
Humans and AI Agents, Working in Parallel
Research · 11 min read ·

Humans and AI Agents, Working in Parallel

In one week, OpenWorker, Buzz, and ego-lite all shipped the same idea: humans and AI agents working in parallel. What the pattern means and where it goes next.

View article: Opus 5 Code Review: Precise but Noisier
Opus 5 Code Review: Precise but Noisier
Research · 8 min read ·

Opus 5 Code Review: Precise but Noisier

In CodeRabbit's own test, Opus 5 wrote more precise review comments but caught fewer known bugs and 4x the nitpicks. What that means for your code review agent.

View article: Opus 5's ARC-AGI-3 Jump Didn't Transfer
Opus 5's ARC-AGI-3 Jump Didn't Transfer
Research · 7 min read ·

Opus 5's ARC-AGI-3 Jump Didn't Transfer

Opus 5's ~4x ARC-AGI-3 lead collapses to a statistical tie with Kimi K3 and Fable 5 on a held-out suite. What the Witness benchmark says about picking a model.

View article: Kimi K3 License: Terms & Commercial Use
Kimi K3 License: Terms & Commercial Use
Research · 9 min read ·

Kimi K3 License: Terms & Commercial Use

Kimi K3 shipped under a bespoke license, not Modified MIT. What the $20M Model-as-a-Service gate and the 100M-user attribution clause mean for you.

View article: Kimi K3 Agent Swarm: 300 Parallel Agents
Kimi K3 Agent Swarm: 300 Parallel Agents
Research · 10 min read ·

Kimi K3 Agent Swarm: 300 Parallel Agents

Kimi K3 Agent Swarm coordinates up to 300 sub-agents and 4,000 tool calls per task. How the architecture works, where it wins, and where it breaks.

View article: AI Agent Video Editors: The MCP Takeover
AI Agent Video Editors: The MCP Takeover
Research · 11 min read ·

AI Agent Video Editors: The MCP Takeover

An AI agent video editor lets you co-edit a timeline by chat, not mouse. See how MCP became the shared interface behind ChatCut, OpenCut, and Palmier Pro.

View article: Sakana Fugu and the End of Model Lock-In
Sakana Fugu and the End of Model Lock-In
Research · 12 min read ·

Sakana Fugu and the End of Model Lock-In

Sakana Fugu routes across frontier LLMs to kill model vendor lock-in. An honest look at the benchmark-vs-reality gap and what it leaves you owning.

View article: OpenSquilla: Token-Efficient AI Agents
OpenSquilla: Token-Efficient AI Agents
Research · 10 min read ·

OpenSquilla: Token-Efficient AI Agents

OpenSquilla is an open-source AI agent runtime that cuts token costs with on-device model routing and layered memory. Here is how it works and who it fits.

View article: What Autonomous AI Agents Should Handle
What Autonomous AI Agents Should Handle
Research · 11 min read ·

What Autonomous AI Agents Should Handle

Autonomous AI agents save time only when the task is clear, bounded, and reviewable. Learn what AI should handle alone and where people stay accountable.