Follow the latest research
alphaXiv connects papers, researchers, and organizations, grounding its answers in the underlying work.
Researchers to follow
View allAre you a researcher? Find your profile
Jointly predicting tokens and multi-token concepts enables language models to reach comparable training loss with substantially fewer tokens than standard architectures.
A shared pixel-space representation lets one model understand, reason about, generate, and edit images while preserving strong multimodal comprehension.
Researchers to follow
View allA frozen video world model can explore new camera paths while preserving an event’s appearance, timing, and previously generated scene states.
The benchmark reveals that multimodal models can recognize topological relations in static scenes but struggle to preserve them while planning actions.