Machine learningLLMs & Text
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Vision and Speech Interaction: A proposed training methodology enables Large Language Models to comprehend visual and speech data, improving speech-to-speech dialogue capabilities and response speed.
Featured in No. 81 on 8 Jan 2025 · 5 days after release · 230 citations today · published in Neural Information Processing Systems
- Released
- 3 Jan 2025
- First featured
- No. 81 · 8 Jan 2025
- Citations (Semantic Scholar)
- 230
- Influential citations
- 26
- Published in
- Neural Information Processing Systems
- Shares when featured
- 52
- Identifier
- arXiv:2501.01957
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).