Papers4

#F1 score

ContextBench: A Benchmark for Context Retrieval in Coding Agents

ContextBench is a new benchmark that checks not just whether a coding AI fixes a bug, but whether it found and used the right pieces of code along the way.

#context retrieval#coding agents#software engineering benchmarks

Not triaged yet

ASA: Training-Free Representation Engineering for Tool-Calling Agents

Intermediate

Youjin Wang, Run Zhou et al.Feb 4arXiv

The paper finds a strange gap: the model’s hidden thoughts almost perfectly show when it should use a tool, but its actual words often don’t trigger the tool under strict rules.

#activation steering#representation engineering#tool calling

Not triaged yet

Ebisu: Benchmarking Large Language Models in Japanese Finance

Intermediate

Xueqing Peng, Ruoyu Xiang et al.Feb 1arXiv

EBISU is a new test that checks how well AI models understand Japanese finance, a language and domain where hints and special terms are common.

#EBISU#Japanese finance NLP#implicit commitment recognition

Not triaged yet

Enhancing Sentiment Classification and Irony Detection in Large Language Models through Advanced Prompt Engineering Techniques

Beginner

Marvin Schmitt, Anne Schwerk et al.Jan 13arXiv

Giving large language models a few good examples and step-by-step instructions can make them much better at spotting feelings in text.

#prompt engineering#few-shot learning#chain-of-thought

Not triaged yet