coder-eval 0.10.2
Evaluate, benchmark, and A/B-test AI coding agents (Claude Code, Codex, Gemini/Antigravity) with sandboxed, reproducible YAML task suites.
Evaluate, benchmark, and A/B-test AI coding agents (Claude Code, Codex, Gemini/Antigravity) with sandboxed, reproducible YAML task suites.

The company has paused training on its next generation of models, called Astra, and its largest planned training run remains on hold, the company said.

IntroductionCitizen science for health invites non-professionals into research, but the degree to which individuals control the knowledge-making process vari...
SAN FRANCISCO, Aug 18 : OpenAI on Tuesday said it is slowing down the pace of its AI model development while it overhauls its research and training systems after OpenAI officials were caught unawares last month when an AI agent under testing hacked another AI…
Python SDK for AgentPub — AI research publication platform