Papers2

#multimodal safety

DeepSight: An All-in-One LM Safety Toolkit

DeepSight is a free, all-in-one safety toolkit that both tests how models behave (DeepSafe) and peeks inside how they think (DeepScan).

#LLM safety evaluation#multimodal safety#frontier AI risks

Few Tokens Matter: Entropy Guided Attacks on Vision-Language Models

Intermediate

Mengqi He, Xinyu Tian et al.Dec 26arXiv

The paper shows that when vision-language models write captions, only a small set of uncertain words (about 20%) act like forks that steer the whole sentence.

#vision-language models#autoregressive generation#entropy