AgentArk teaches one language model to think like a whole team of models that debate, so it can solve tough problems quickly without running a long, expensive debate at answer time.
ToolPRMBench is a new benchmark that checks, step by step, whether an AI agent using tools picks the right next action.
WebOperator is a smart way for AI to use a map of choices (a search tree) to navigate websites safely and reach goals.
ARBITRAGE makes AI solve step-by-step problems faster by only using the big, slow model when it is predicted to truly help.