How I Study AI - Learn AI Papers & Lectures the Easy Way

Locality-Attending Vision Transformer

Sina Hajimiri, Farzad Beizaee et al.Mar 5arXiv

Vision Transformers (ViTs) are great at recognizing what is in a whole image but often blur the tiny details needed to label each pixel (segmentation).

#Vision Transformer#self-attention#segmentation

Not triaged yet

A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning

Intermediate

Zixin Zhang, Kanghao Chen et al.Dec 16arXiv

This paper builds A4-Agent, a smart three-part helper that figures out where to touch or use an object just from a picture and a written instruction, without any extra training.

#affordance prediction#zero-shot learning#vision-language models

Not triaged yet

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation

Beginner

Jiehui Huang, Yuechen Zhang et al.Dec 8arXiv

UnityVideo is a single, unified model that learns from many kinds of video information at once—like colors (RGB), depth, motion (optical flow), body pose, skeletons, and segmentation—to make smarter, more realistic videos.

#multimodal video generation#multi-task learning#dynamic noise scheduling

Not triaged yet

Papers3

Locality-Attending Vision Transformer

A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation