Vai al contenuto principale
Torna al leaderboard

business

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

📄 Paper Trending (146⬆️): Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, quer...

Fonte: HuggingFace_Papers

Cosa fa

📄 Paper Trending (146⬆️): Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, quer...

Per chi è

Non disponibile

Prezzo

Non disponibile

Punti di forza

  • Non disponibile

Limiti

  • Non disponibile

Vuoi integrarlo nel tuo business? Prenota una call gratuita

Prenota ora
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs | ToolAI - StudioCentOS