ReToken: Visual Retrieval with Qwen3-VL-8B
This demo showcases ReToken-Qwen3VL-8B, a vision-language model augmented with a learned retrieval token for improved visual retrieval. Ask questions about images — the model leverages a fine-tuned Qwen3-VL-8B backbone trained with a retrieval token that improves long-context visual understanding.
Examples