A terminal AI agent that works on a codebase, container, browser or desktop, uses what it built, and judges its work against criteria written before...
A terminal AI agent that works on a codebase, container, browser or desktop, uses what it built, and judges its work against criteria written before it started. One agent loop, rubric-first evaluator, tiered memory, four-vendor routing, six benchmark harnesses.
Official website restored from the pre-incident audit; product details pending editorial verification.
LMSYS Chatbot Arena is a crowdsourced open platform for LLM evals. Collected over 1,000,000 human pairwise comparisons to rank LLMs with the Bradley-Terry model and display the model ratings in Elo-scale.
清华技术AI对话助手
字节跳动AI助手