Researchers in the lab of Yisong Yue, a professor of computing and mathematical sciences at Caltech, are building efficient multimodal foundation models: AI systems that learn to understand video the way people do. In partnership with The Bright Initiative, the lab is developing models that train far more efficiently than frontier closed-source models, and releasing the results openly so that the wider research community can build on them.
The models are aimed at keeping frontier video understanding within the reach of academic and public-interest research, and not just industry laboratories. Because the models’ training procedures and checkpoints are openly released, they are available for use by other research groups.
Challenge
Video recordings now document most of our lives, and also can be vital to scientific research. A model that can “reason” and interpret the content in videos can enable important downstream applications in domains ranging from robotics to environmental monitoring, and lead to scientific innovation. However, most frontier video foundation models are closed: their training data and weights are not published. A researcher who wants to apply one of these models to a clinical recording, a robot’s camera feed, or footage of a changing landscape cannot adapt it to their own domain, and often cannot afford to run it at the scale their question demands. The Caltech team set out to close that gap with an efficient, openly released model that reads video in far finer detail.
Impact
Training on a broader and more varied body of public video measurably improved the model’s video understanding. The lab is now extending the work to long-form footage. Dense video understanding has developed largely inside private laboratories; this work aims to provide a shared video foundation, so that researchers across the scientific community can adapt these capabilities to their own questions and build the applications that matter in their domains.