Machine learningLLMs & Text
Ferret: Refer and Ground Anything Anywhere at Any Granularity
Spatial Referring in Images: Ferret is a Multimodal Large Language Model that can understand and locate spatial references in images, outperforming other models in region-based and localization-required multimodal chatting.
Featured in No. 21 on 16 Oct 2023 · 5 days after release · 604 citations today · published in International Conference on Learning Representations
- Released
- 11 Oct 2023
- First featured
- No. 21 · 16 Oct 2023
- Citations (Semantic Scholar)
- 604
- Influential citations
- 71
- Published in
- International Conference on Learning Representations
- Shares when featured
- 52
- Identifier
- arXiv:2310.07704
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).