Instructions to use BrightGuo/ProcObject-Qwen3-VL-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BrightGuo/ProcObject-Qwen3-VL-4B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("BrightGuo/ProcObject-Qwen3-VL-4B") model = AutoModelForMultimodalLM.from_pretrained("BrightGuo/ProcObject-Qwen3-VL-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
ProcObject-Qwen3-VL-4B
Qwen3-VL-4B-Instruct fine-tuned with object-centric SFT on ProcObject-10K — the fine-tuned model of ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos (NeurIPS 2026, Evaluations & Datasets Track). Paper · Code · Dataset
Given frames of an instructional video clip and a question about an object, the model answers and localizes the supporting moments:
{"answer": "The tortilla starts whole on the table, is torn into two pieces, and finally placed into a bowl.", "evidence": [[0, 8], [14, 16]]}
evidence intervals are in seconds from the start of the clip.
Training
Two stages on the 9,472 ProcObject-10K training QA (8 GPUs, bf16):
- Evidence-prediction SFT — LoRA (r=16, α=32) on the language model's linear layers, trained on the JSON answer + evidence target; 3 epochs, lr 5e-5; 2 FPS, ≤48 frames.
- Object-centric SFT — continues the same LoRA, adds LoRA on the attention layers of the top
third of the vision blocks, tunes the vision merger, and trains two auxiliary heads: a spatial
head supervised by soft patch masks from Grounding DINO boxes of LLM-extracted object phrases,
and a temporal head supervised by frame-in-evidence labels.
L = L_gen + 0.05·L_spl + 0.10·L_tmp, 3 epochs.
The auxiliary heads are training-only and are not part of this checkpoint: it is a standard merged Qwen3-VL model with no extra inference cost. Full recipe and code: github.com/WenliangGuo/ProcObject-10K.
Results on the ProcObject-10K test set
With the released evaluation harness (48 frames; J. = 0–5 LLM judge, mean of Qwen3-4B and Llama-3.2-3B):
| S. | B. | J. | mIoU | mIoP | mIoG | R@0.3 | |
|---|---|---|---|---|---|---|---|
| Qwen3-VL-4B-Instruct (zero-shot, paper) | 73.3 | 89.2 | – | 38.6 | 63.1 | 54.1 | – |
| this model | 80.8 | 92.0 | 3.25 | 45.3 | 66.8 | 56.4 | 63.5 |
mIoU is 51.9 on Multi-hop Reasoning and 35.9 on Needle-in-a-Haystack questions.
Usage
The reported numbers use the benchmark protocol: 48 frames sampled uniformly from the clip, resized
to 512×512, each sent as an image preceded by [Timestamp: <t>s], after the system prompt in
benchmark/prompts/system_prompt.txt,
greedy decoding. The easiest way to reproduce them is the repository's harness:
git clone https://github.com/WenliangGuo/ProcObject-10K && cd ProcObject-10K
pip install -r requirements/eval.txt
# build data/clips/ first (preprocess/README.md), then:
MODEL=BrightGuo/ProcObject-Qwen3-VL-4B NAME=procobject_qwen3vl_4b NUM_FRAMES=48 \
bash benchmark/scripts/run_sharded.sh 0
bash benchmark/scripts/evaluate.sh benchmark/pred_results/procobject_qwen3vl_4b_predictions.json
Or serve it with vLLM (vllm serve BrightGuo/ProcObject-Qwen3-VL-4B --limit-mm-per-prompt '{"image": 64}')
and send the same message layout through the OpenAI-compatible API (benchmark/models/runners.py,
VLLMRunner).
Tested with vLLM 0.17.1 / transformers 4.57.6.
License
CC BY-NC 4.0, following the ProcObject-10K annotations it was trained on; the base model Qwen3-VL-4B-Instruct is Apache-2.0.
Citation
@inproceedings{guo2026procobject,
title = {{ProcObject-10K}: Benchmarking Object-Centric Procedural Understanding in Instructional Videos},
author = {Guo, Wenliang and Kong, Yu},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track},
year = {2026},
url = {https://arxiv.org/abs/2512.03479}
}
- Downloads last month
- 21