Papers4

All Beginner Intermediate Advanced

All Sources arXiv

#orchestrator

General Agent Evaluation

Intermediate

Elron Bandel, Asaf Yehudai et al.Feb 26arXiv

This paper shows how to fairly test "general-purpose" AI agents that should work in many places without special tweaks.

#general-purpose agents#agent evaluation#unified protocol

Not triaged yet

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

Intermediate

Bowen Yang, Kaiming Jin et al.Jan 12arXiv

Computer-using agents kept forgetting important visual details over long tasks and could not reliably find up-to-date, step-by-step help for unfamiliar apps.

#computer-using agents#vision-language models#milestone memory

Not triaged yet

OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs

Beginner

Xin Wang, Yunhao Chen et al.Jan 4arXiv

OpenRT is a big, open-source test bench that safely stress-tests AI models that handle both text and images.

#OpenRT#red teaming#multimodal LLM

Not triaged yet

Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

Beginner

Sherman Wong, Zhenting Qi et al.Dec 11arXiv

This paper introduces the Confucius Code Agent (CCA), a coding helper built to handle huge real-world codebases with long tasks and many tools.

#coding agents#agent scaffolding#context management

Not triaged yet