Papers1055

INTELLECT-3: Technical Report

Prime Intellect Team, Mika Senghaas et al.Dec 18arXiv

INTELLECT-3 is a 106B-parameter Mixture-of-Experts model (about 12B active per token) trained with large-scale reinforcement learning and it beats many bigger models on math, coding, science, and reasoning tests.

#INTELLECT-3#prime-rl#verifiers

Not triaged yet

ModelTables: A Corpus of Tables about Models

Intermediate

Zhengyuan Dong, Victor Zhong et al.Dec 18arXiv

ModelTables is a giant, organized collection of tables that describe AI models, gathered from Hugging Face model cards, GitHub READMEs, and research papers.

#Model Lake#Model Cards#Scientific Tables

Not triaged yet

TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times

Intermediate

Jintao Zhang, Kaiwen Zheng et al.Dec 18arXiv

TurboDiffusion speeds up video diffusion models by about 100–200 times while keeping video quality comparable.

#TurboDiffusion#video diffusion acceleration#Sparse-Linear Attention

Not triaged yet

Are We on the Right Way to Assessing LLM-as-a-Judge?

Intermediate

Yuanning Feng, Sinan Wang et al.Dec 17arXiv

This paper asks whether we are judging AI answers the right way and introduces Sage, a new way to test AI judges without using human-graded answers.

#LLM-as-a-Judge#Sage evaluation#Intra-Pair Instability

Not triaged yet

Spatia: Video Generation with Updatable Spatial Memory

Intermediate

Jinjing Zhao, Fangyun Wei et al.Dec 17arXiv

Spatia is a video generator that keeps a live 3D map of the scene (a point cloud) as its memory while making videos.

#video generation#spatial memory#3D point cloud

Not triaged yet

In Pursuit of Pixel Supervision for Visual Pre-training

Intermediate

Lihe Yang, Shang-Wen Li et al.Dec 17arXiv

Pixels are the raw stuff of images, and this paper shows you can learn great vision skills by predicting pixels directly, not by comparing fancy hidden features.

#pixel supervision#masked autoencoders#MAE redesign

Not triaged yet

DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

Intermediate

Lunbin Zeng, Jingfeng Yao et al.Dec 17arXiv

This paper shows a simple way to turn any strong autoregressive (step-by-step) model into a diffusion vision-language model (parallel, block-by-block) without changing the architecture.

#DiffusionVL#diffusion vision-language model#block diffusion

Not triaged yet

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

Intermediate

Yuwei Guo, Ceyuan Yang et al.Dec 17arXiv

This paper fixes a common problem in video-making AIs where tiny mistakes snowball over time and ruin long videos.

#autoregressive video diffusion#exposure bias#teacher forcing

Not triaged yet

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning

Intermediate

Yifei Li, Wenzhao Zheng et al.Dec 17arXiv

Skyra is a detective-style AI that spots tiny visual mistakes (artifacts) in videos to tell if they are real or AI-generated, and it explains its decision with times and places in the video.

#AI-generated video detection#artifact reasoning#multimodal large language model

Not triaged yet

Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning

Intermediate

Zhenwen Liang, Sidi Lu et al.Dec 17arXiv

This paper teaches large language models (LLMs) to explore smarter by listening to their own gradients—the directions they would update—rather than chasing random variety.

#gradient-guided reinforcement learning#GRL#GRPO

Not triaged yet

VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?

Intermediate

Hongbo Zhao, Meng Wang et al.Dec 17arXiv

Long texts are expensive for AI to read because each extra token costs a lot of compute and memory.

#vision‑text compression#VTCBench#vision‑language models

Not triaged yet

IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning

Intermediate

Yuanhang Li, Yiren Song et al.Dec 17arXiv

IC-Effect is a new way to add special effects to existing videos by following a text instruction while keeping everything else unchanged.

#video editing#visual effects#diffusion transformer

Not triaged yet

73 74 75 76 77