NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
IntermediateJingzhe Ding, Shengda Long et al.Dec 14arXiv
NL2Repo-Bench is a new benchmark that tests if coding agents can build a whole Python library from just one long natural-language document and an empty folder.
#NL2Repo-Bench#autonomous coding agents#long-horizon reasoning