Skip to main content
Back to leaderboard

llm

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

📄 Trending Paper (47⬆️): Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning ofte...

Source: HuggingFace_Papers

What it does

📄 Trending Paper (47⬆️): Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning ofte...

Who is it for

Not available

Pricing

Not available

Strengths

  • Not available

Limits

  • Not available

Want to integrate it in your business? Book a free call

Book now
ReferTrack: Referring Then Tracking for Embodied Visual Tracking | ToolAI - StudioCentOS