multimodal
fact
bullish
A self-supervised framework can project four modalities into a shared 256-dimensional embedding space and discover aesthetic structure through iterative clustering
We present a self-supervised framework that projects four modalities (text, audio, image and video) into a shared 256-dimensional embedding space and applies iterative clustering to discover aesthetic structure.
Machine Learning30 Aug 2026