ReToken: Visual Retrieval with Qwen3-VL-8B

This demo showcases ReToken-Qwen3VL-8B, a vision-language model augmented with a learned retrieval token for improved visual retrieval. Ask questions about images — the model leverages a fine-tuned Qwen3-VL-8B backbone trained with a retrieval token that improves long-context visual understanding.

Paper · GitHub · Model

Examples