LOCA-bench is a test that challenges AI agents to work correctly as their to-do list and background information grow very, very long.
AgentLongBench is a new test that checks how well AI agents think over very long stories made of their own actions and the world's replies, not just by reading static documents.