Object Detection · Retina

What is Qwen?

See how Qwen-VL turns image patches and a prompt into language and object locations. Learn what the position-aware adapter does, why grounding is generated token by token, and when a general vision-language model is useful in robotics.

Imagine asking a warehouse robot to find “the damaged carton beside the blue pallet.” Qwen can connect that description to a region in the camera image, allowing the robot to look for more than a short, fixed list of object names.

1. What is Qwen?

A warehouse instruction often contains more than an object name. “Pick the damaged carton beside the blue pallet” includes an object, a property, and a relationship. A fixed detector may recognize carton, but it was not necessarily trained to reason about damaged, blue, or beside.

Qwen-VL approaches the problem as a vision-language model. It receives both the camera image and the instruction, then generates a response informed by both. When the response connects a phrase to an image region, the task is called visual grounding.

Grounding is more demanding than classification. Classification asks, “Is there a carton?” Grounding asks, “Which pixels correspond to the carton described by this particular phrase?” The surrounding words help distinguish the intended carton from other cartons in the same scene.

Qwen-VL is a generalist rather than a dedicated detector. The same model can describe an image, answer a question, read visible text, continue a dialogue, or produce a location. This breadth gives it semantic flexibility. It also means that localization is represented as generated language and coordinates, rather than through a fixed detection head built only for boxes.

1.1 See the Full Qwen-VL Pipeline

Image + prompt → grounded response

ImageVisual evidence
Text promptRequested objects
Vision TransformerPatch-level visual tokens
Position-aware adapterCompresses to 256 tokens
Qwen language modelAutoregressive decoding
Grounded responseLanguage and locations

Position-aware visual tokens join the same generated sequence as language.

Qwen-VL compresses spatial visual features before combining them with language for grounded generation. Adapted as an explanatory diagram fromBai et al., “Qwen-VL”.

Read the diagram from top to bottom. Image and Text prompt enter together. The image passes through the Vision Transformer and Position-aware adapter before reaching the Qwen language model. The model combines those visual tokens with the prompt and produces a Grounded response containing language and locations.

1.2 Follow an Image and Prompt Through the Model

Follow the instruction the damaged carton beside the blue pallet through the model.

The diagram begins at Image. In the original Qwen-VL paper, the Vision Transformer acts as the visual receptor. It divides the image into patches and converts them into feature vectors. Each vector describes part of the image, but a high-resolution image can produce a long sequence of such features.

Next comes the Position-aware adapter. It creates a bridge between the vision model and the language model. A single cross-attention layer uses learned queries to gather the relevant visual information and compress the variable image sequence into 256 visual tokens. This step matters because the language model needs a compact visual representation without losing the spatial clues required for localization.

Finally, those visual tokens and the Text prompt enter the Qwen language model—Qwen-7B in the original paper. The language model processes them together and generates a response one token at a time. It can use words such as damaged and beside because those words participate in the same context as the image representation.

Boxes are made part of this generated language. The paper normalizes image coordinates to a 0–999 range and places them between special region tokens. During training, Qwen-VL sees examples that connect images, captions, and boxes. It gradually learns that a phrase can be followed by coordinates describing the corresponding region.

The final box in the diagram is the Grounded response. Some generated tokens describe the object and others encode its location. There is no separate fixed-class detection head at the end.

1.3 How Qwen-VL Learns Language and Location

Qwen-VL’s central innovation is to make localization part of a general multimodal conversation. The position-aware adapter compresses image features while retaining the spatial information needed to point back into the image. The unified input–output format lets words, images, and boxes participate in one sequence.

The model learns this ability in stages. Large-scale image–text pre-training first teaches broad visual concepts while the language model remains frozen. Multi-task pre-training then uses higher-resolution images and trains all components together on a wider mixture of tasks, including grounding and text reading. Supervised instruction tuning finally produces Qwen-VL-Chat, which learns to follow conversational instructions.

The order matters. The model first learns what visual content looks like, then learns to solve specific multimodal tasks, and finally learns to respond helpfully to instructions. Grounding is therefore connected to the model’s wider language and visual knowledge rather than trained as an isolated output.

1.4 Why Qwen-VL Generates One Token at a Time

Because the response is generated as a sequence, each new token depends on the image, the prompt, and the tokens already produced. Let II be the image, pp the text prompt, and yky_k the token generated at step kk. The conditional output distribution is

P(y∣I,p)=∏k=1TP(yk∣I,p,y1,…,yk−1).P(\mathbf{y}\mid I,p)=\prod_{k=1}^{T}P(y_k\mid I,p,y_1,\ldots,y_{k-1}).

The equation says that the model builds its answer step by step. Some tokens describe the object; others encode its location. A dedicated detector instead predicts class scores and box vectors directly and in parallel.

This difference explains the practical trade-offs:

  • Wording can change localization because text is part of the conditioning context.
  • The model can use broad semantic knowledge unavailable to a small fixed-class head.
  • Output validation is essential because generative models can produce plausible but unsupported answers.
  • Latency and determinism may differ from specialized detectors.

For a robot, Qwen is most useful when language carries information that a fixed class label would lose. Its output should still be treated as a hypothesis. Before acting, the system must validate the returned region, connect it to depth or segmentation, and check whether the proposed action is geometrically and physically safe.

2. How to Use Qwen in the Telekinesis Agentic OS

Telekinesis provides Qwen grounding as the Retina Skill detect_objects_using_qwen. The code below loads an image, requests three object names, and prints the normalized categories, scores, and boxes returned by the Skill.

from telekinesis import datatypes, retina

image = datatypes.Image.from_url(
    url="https://assets.telekinesis.ai/examples/v1/images/warehouse_1.jpg"
)

detections, categories = retina.detect_objects_using_qwen(
    image=image,
    objects=["person", "forklift", "pallet"],
)

for detection in detections:
    label = categories[detection.category_id]
    print(label.name, detection.score, detection.bbox)

The visualization below shows the grounded output rather than free-form model prose. Inspect whether every requested concept is supported by visible evidence, especially on negative scenes and ambiguous prompts.

Qwen object detection output in a warehouse scene

Qwen uses multimodal understanding to localize requested names. Treat the boxes as hypotheses to validate against task geometry and negative scenes.

Try it out

Build with Qwen in Retina

See the complete Skill signature, input types, output schema, and current examples.

Open Qwen docs

Runnable example

Run the Qwen detection example

Open the exact Python script used to detect requested objects with Qwen and Retina.

View Python example

3. Benchmarking

3.1 Results Reported in the Qwen-VL Paper

The Qwen-VL paper evaluates referring-expression comprehension as localization accuracy. The abbreviated table below reports selected validation/test results from Table 6; it describes the original Qwen-VL checkpoints, not necessarily the current hosted service version.

ModelRefCOCO valRefCOCO test-ARefCOCO test-BRefCOCO+ valRefCOCOg val
Shikra-13B87.8391.1181.8182.8982.64
Qwen-VL-7B89.3692.2685.3483.1285.58
Qwen-VL-7B-Chat88.5592.2784.5182.8285.96
Grounding DINO-L90.5693.1988.2482.7586.13

RefCOCO results cannot predict warehouse performance by themselves. Build a prompt-by-scene test set and report localization success at an IoU tolerance, hallucination rate on negative images, end-to-end latency, and final robot-task success.

3.2 Comparison with Other Telekinesis Skills

SkillQueryThreshold controlStrongest reason to choose it
QwenList of object namesNo exposed tuning thresholdSemantic and multimodal flexibility
Grounding DINOPhrase promptBox and text thresholdsDirect open-vocabulary grounding control
RF-DETRCOCO class setScore thresholdEnd-to-end fixed-vocabulary detector
YOLOXCOCO class setScore and NMS thresholdsEfficient dense detector with explicit NMS

4. Where to Go Next?

Study Grounding DINO for an open-vocabulary detector with separate localization and text-alignment thresholds. For fixed COCO classes, compare RF-DETR and YOLOX.

5. References