Vision-R1
[ICLR2026] This is the first paper to explore how to effectively use R1-like RL for MLLMs and introduce Vision-R1, a reasoning MLLM that leverages cold-start initialization and RL training to incentivize reasoning capability.
파일 탐색기
최종 버전 다운로드 (.zip)- AMNH.jpg
- dog.jpg
- running.jpg
- tesla.jpg
- tokyo.jpg
- example.tsv
- image_queries.jsonl
- example.tsv
- example_dev.csv
- example_val.csv
- example.jsonl
- test.tsv
- corpus.jsonl
- queries.jsonl
- dockerfile
- dingding.png
- discord_qr.jpg
- wechat.png
- evalscope_framework.png
- evalscope_icon.png
- evalscope_icon.svg
- evalscope_icon_dark.png
- evalscope_logo.png
- evaluate.md
- index.md
- sample.md
- schema.md
- clip.md
- embedding.md
- index.md
- llm.md
- vlm.md
- mmlu_pro_preview.png
- add_benchmark.md
- custom_model.md
- llm_full_stack.md
- swift_integration.md
- MM_RAG_LangChain_1.png
- RAG_Pipeline_1.png
- RAG_Evaluation.md
- index.md
- index.md
- mmlu.md
- index.md
- QwQ-32B-Preview.md
- collection.png
- model_compare.png
- model_prediction.png
- report_details.png
- report_overview.png
- setting.png
- single_dataset.png
- single_model.png
- basic_usage.md
- installation.md
- introduction.md
- parameters.md
- supported_dataset.md
- visualization.md
- longwriter.md
- toolbench.md
- eval_result.png
- generation_process.png
- clip_benchmark.md
- index.md
- mteb.md
- ragas.md
- index.md
- opencompass_backend.md
- vlmevalkit_backend.md
- qwen_speed_benchmark.png
- custom.md
- examples.md
- index.md
- parameters.md
- quick_start.md
- speed_benchmark.md
- arena.md
- .readthedocs.yaml
- conf.py
- index.md
- Makefile
- evalscope_framework.png
- evalscope_icon.png
- evalscope_icon.svg
- evalscope_icon_dark.png
- evalscope_logo.png
- evaluate.md
- index.md
- sample.md
- schema.md
- clip.md
- embedding.md
- index.md
- llm.md
- vlm.md
- mmlu_pro_preview.png
- add_benchmark.md
- custom_model.md
- deepseek_r1_distill.md
- llm_full_stack.md
- swift_integration.md
- AIME.png
- claude3_5.png
- codeforces.png
- human_preference.png
- MMLU.png
- o1_preview.png
- evaluation.md
- 1e548e9b-2e72-4cae-bfb7-d5e836bccbc7.png
- 284d1152-8aa9-4b75-ae31-d2d5bcea6421.png
- 2d029f6f-b558-48a2-9f07-e7d0e51ac77a.png
- 3ab7bf8f-c524-468c-9571-5b7822968795.png
- 3cedcbd3-cc9f-4f04-9a60-14b10ad48498.png
- 4e9a558a-c91c-469a-9e03-3c3817357c26.png
- 8d90a034-04c6-4102-85e1-a4b43d47212e.png
- b2115395-9bf1-4bfd-92de-443ddd88611d.png
- c87a577f-9d42-44dd-a7e8-3e7b524b4133.png
- c8d41922-f821-46db-a908-76d644f1e5e5.jpg
- cebbcd76-5546-4f0f-8634-920a996f8b0c.png
- ec6afe12-d723-4539-90bc-5c705cb4cbb6.png
- f2754af7-d09b-47f4-8525-7ef478fb8f56.jpg
- f8dcf95b-de44-45af-a96d-65f6466cd39c.png
- MM_RAG_LangChain_1.png
- rag_pipeline.png
- RAG_Pipeline_1.png
- download_images.py
- multimodal_RAG.md
- RAG_Evaluation.md
- index.md
- index.md
- mmlu.md
- index.md
- QwQ-32B-Preview.md
- collection.png
- model_compare.png
- model_prediction.png
- report_details.png
- report_overview.png
- setting.png
- single_dataset.png
- single_model.png
- basic_usage.md
- installation.md
- introduction.md
- parameters.md
- supported_dataset.md
- visualization.md
- longwriter.md
- toolbench.md
- eval_result.png
- generation_process.png
- clip_benchmark.md
- index.md
- mteb.md
- ragas.md
- index.md
- opencompass_backend.md
- vlmevalkit_backend.md
- qwen_speed_benchmark.png
- custom.md
- examples.md
- index.md
- parameters.md
- quick_start.md
- speed_benchmark.md
- arena.md
- .readthedocs.yaml
- conf.py
- index.md
- Makefile
- __init__.py
- eval_api.py
- eval_datasets.py
- __init__.py
- api_meta_template.py
- backend_manager.py
- __init__.py
- image_caption.py
- zeroshot_classification.py
- zeroshot_retrieval.py
- webdataset_convert.py
- webdatasets.txt
- __init__.py
- arguments.py
- dataset_builder.py
- task_template.py
- __init__.py
- Classification.py
- Clustering.py
- CustomTask.py
- PairClassification.py
- Reranking.py
- Retrieval.py
- STS.py
- __init__.py
- arguments.py
- base.py
- task_template.py
- persona_prompt.py
- __init__.py
- build_distribution.py
- build_transform.py
- testset_generation.py
- translate_prompt.py
- __init__.py
- arguments.py
- task_template.py
- __init__.py
- clip.py
- embedding.py
- llm.py
- tools.py
- __init__.py
- backend_manager.py
- __init__.py
- backend_manager.py
- custom_dataset.py
- __init__.py
- base.py
- __init__.py
- aime24_adapter.py
- aime25_adapter.py
- __init__.py
- ai2_arc.py
- arc_adapter.py
- boolean_expressions.txt
- causal_judgement.txt
- date_understanding.txt
- disambiguation_qa.txt
- dyck_languages.txt
- formal_fallacies.txt
- geometric_shapes.txt
- hyperbaton.txt
- logical_deduction_five_objects.txt
- logical_deduction_seven_objects.txt
- logical_deduction_three_objects.txt
- movie_recommendation.txt
- multistep_arithmetic_two.txt
- navigate.txt
- object_counting.txt
- penguins_in_a_table.txt
- reasoning_about_colored_objects.txt
- ruin_names.txt
- salient_translation_error_detection.txt
- snarks.txt
- sports_understanding.txt
- temporal_sequences.txt
- tracking_shuffled_objects_five_objects.txt
- tracking_shuffled_objects_seven_objects.txt
- tracking_shuffled_objects_three_objects.txt
- web_of_lies.txt
- word_sorting.txt
- __init__.py
- bbh_adapter.py
- __init__.py
- ceval_adapter.py
- ceval_exam.py
- __init__.py
- cmmlu.py
- cmmlu_adapter.py
- samples.jsonl
- __init__.py
- competition_math.py
- competition_math_adapter.py
- __init__.py
- data_collection_adapter.py
- __init__.py
- general_mcq_adapter.py
- __init__.py
- general_qa_adapter.py
- __init__.py
- chain_of_thought.txt
- gpqa_adapter.py
- __init__.py
- gsm8k.py
- gsm8k_adapter.py
- __init__.py
- hellaswag.py
- hellaswag_adapter.py
- __init__.py
- humaneval.py
- humaneval_adapter.py
- __init__.py
- ifeval_adapter.py
- instructions.py
- instructions_registry.py
- instructions_util.py
- utils.py
- __init__.py
- iquiz_adapter.py
- __init__.py
- math_500_adapter.py
- __init__.py
- mmlu.py
- mmlu_adapter.py
- samples.jsonl
- __init__.py
- mmlu_pro_adapter.py
- __init__.py
- race.py
- race_adapter.py
- samples.jsonl
- __init__.py
- samples.jsonl
- trivia_qa.py
- trivia_qa_adapter.py
- __init__.py
- truthful_qa.py
- truthful_qa_adapter.py
- __init__.py
- benchmark.py
- data_adapter.py
- __init__.py
- base.py
- cli.py
- start_app.py
- start_eval.py
- start_perf.py
- start_server.py
- __init__.py
- evaluator.py
- sampler.py
- schema.py
- __init__.py
- auto_reviewer.py
- __init__.py
- evaluator.py
- rating_eval.py
- __init__.py
- rouge_scorer.py
- gpt2-zhcn3-v4.bpe
- gpt2-zhcn3-v4.json
- __init__.py
- code_metric.py
- math_parser.py
- metrics.py
- named_metrics.py
- rouge_metric.py
- __init__.py
- custom_model.py
- dummy_model.py
- __init__.py
- base_adapter.py
- chat_adapter.py
- choice_adapter.py
- custom_adapter.py
- local_model.py
- model.py
- server_adapter.py
- __init__.py
- base.py
- custom_api.py
- dashscope_api.py
- openai_api.py
- __init__.py
- base.py
- custom.py
- flickr8k.py
- line_by_line.py
- longalpaca.py
- openqa.py
- speed_benchmark.py
- __init__.py
- registry.py
- __init__.py
- analysis_result.py
- benchmark_util.py
- db_util.py
- handler.py
- local_server.py
- __init__.py
- arguments.py
- benchmark.py
- http_client.py
- main.py
- cfg_arena.yaml
- cfg_arena_zhihu.yaml
- cfg_pairwise_baseline.yaml
- cfg_single.yaml
- lmsys_v2.jsonl
- prompt_templates.jsonl
- battle.jsonl
- category_mapping.yaml
- question.jsonl
- arc.yaml
- bbh.yaml
- bbh_mini.yaml
- ceval.yaml
- ceval_mini.yaml
- cmmlu.yaml
- eval_qwen-7b-chat_v100.yaml
- general_qa.yaml
- gsm8k.yaml
- mmlu.yaml
- mmlu_mini.yaml
- __init__.py
- __init__.py
- app.py
- combinator.py
- generator.py
- utils.py
- __init__.py
- judge.txt
- longbench_write.jsonl
- longbench_write_en.jsonl
- longwrite_ruler.jsonl
- __init__.py
- data_etl.py
- openai_api.py
- __init__.py
- default_task.json
- default_task.yaml
- eval.py
- infer.py
- longbench_write.py
- README.md
- utils.py
- __init__.py
- swift_infer.py
- __init__.py
- config_default.json
- config_default.yaml
- eval.py
- infer.py
- README.md
- requirements.txt
- toolbench_static.py
- __init__.py
- __init__.py
- arena_utils.py
- chat_service.py
- completion_parsers.py
- io_utils.py
- logger.py
- model_utils.py
- utils.py
- __init__.py
- arguments.py
- config.py
- constants.py
- run.py
- run_arena.py
- summarizer.py
- version.py
- default_eval_swift_openai_api.json
- default_eval_swift_openai_api.yaml
- eval_qwen_oc_cfg.yaml
- eval_vlm_local.yaml
- eval_vlm_swift.yaml
- task_config_8fafb3.yaml
- arc_ARC-Challenge.jsonl
- arc_ARC-Easy.jsonl
- ceval_college_programming.jsonl
- ceval_computer_architecture.jsonl
- ceval_computer_network.jsonl
- ceval_operating_system.jsonl
- gsm8k_main.jsonl
- humaneval_openai_humaneval.jsonl
- ifeval_default.jsonl
- arc.json
- ceval.json
- gsm8k.json
- humaneval.json
- ifeval.json
- arc_ARC-Challenge.jsonl
- arc_ARC-Easy.jsonl
- ceval_college_programming.jsonl
- ceval_computer_architecture.jsonl
- ceval_computer_network.jsonl
- ceval_operating_system.jsonl
- gsm8k_main.jsonl
- humaneval_openai_humaneval.jsonl
- ifeval_default.jsonl
- task_config_46cf67.yaml
- arc_ARC-Challenge.jsonl
- arc_ARC-Easy.jsonl
- ceval_college_programming.jsonl
- ceval_computer_architecture.jsonl
- ceval_computer_network.jsonl
- ceval_operating_system.jsonl
- gsm8k_main.jsonl
- humaneval_openai_humaneval.jsonl
- ifeval_default.jsonl
- arc.json
- ceval.json
- gsm8k.json
- humaneval.json
- ifeval.json
- arc_ARC-Challenge.jsonl
- arc_ARC-Easy.jsonl
- ceval_college_programming.jsonl
- ceval_computer_architecture.jsonl
- ceval_computer_network.jsonl
- ceval_operating_system.jsonl
- gsm8k_main.jsonl
- humaneval_openai_humaneval.jsonl
- ifeval_default.jsonl
- example_deepseek_distill_collection.py
- example_eval_clip_benchmark.py
- example_eval_custom_llm_data.py
- example_eval_mteb.py
- example_eval_perf.py
- example_eval_rag.py
- example_eval_swift_openai_api.py
- example_eval_toolbench.py
- example_eval_vlm_local.py
- example_eval_vlm_swift.py
- example_swift_eval.py
- app.txt
- docs.txt
- framework.txt
- inner.txt
- opencompass.txt
- perf.txt
- rag.txt
- tests.txt
- vlmeval.txt
- run_all.py
- __init__.py
- test_collection.py
- test_run.py
- __init__.py
- test_perf.py
- __init__.py
- test_clip_benchmark.py
- test_mteb.py
- test_ragas.py
- __init__.py
- test_run_swift_eval.py
- test_run_swift_vlm_eval.py
- test_run_swift_vlm_jugde_eval.py
- __init__.py
- test_vlmeval.py
- __init__.py
- test_run_all.py
- __init__.py
- bailingmm.py
- base.py
- bluelm_v_api.py
- claude.py
- cloudwalk.py
- doubao_vl_api.py
- gemini.py
- glm_vision.py
- gpt.py
- hf_chat_model.py
- hunyuan.py
- jt_vl_chat.py
- lmdeploy.py
- qwen_api.py
- qwen_vl_api.py
- reka.py
- sensechat_vision.py
- siliconflow.py
- stepai.py
- taichu.py
- taiyi.py
- __init__.py
- common.py
- doc_parsing_evaluator.py
- kie_evaluator.py
- ocr_evaluator.py
- __init__.py
- cgbench.py
- crpe.py
- hrbench.py
- judge_util.py
- llavabench.py
- logicvista.py
- longvideobench.py
- mathv.py
- mathverse.py
- mathvista.py
- mlvu.py
- mmbench_video.py
- mmdu.py
- mmniah.py
- mmvet.py
- multiple_choice.py
- mvbench.py
- naturalbench.py
- ocrbench.py
- olympiadbench.py
- qspatial.py
- tablevqabench.py
- tempcompass.py
- videomme.py
- vqa_eval.py
- wemath.py
- yorn.py
- __init__.py
- cgbench.py
- cmmmu.py
- dude.py
- dynamath.py
- image_base.py
- image_caption.py
- image_ccocr.py
- image_mcq.py
- image_mt.py
- image_vqa.py
- image_yorn.py
- longvideobench.py
- miabench.py
- mlvu.py
- mmbench_video.py
- mmgenbench.py
- mmlongbench.py
- mmmath.py
- mvbench.py
- slidevqa.py
- tempcompass.py
- text_base.py
- text_mcq.py
- vcr.py
- video_base.py
- video_concat_dataset.py
- video_dataset_config.py
- videomme.py
- vl_rewardbench.py
- wildvision.py
- __init__.py
- file.py
- log.py
- misc.py
- vlm.py
- __init__.py
- arguments.py
- matching_util.py
- mp_util.py
- result_transfer.py
- __init__.py
- internvl_chat.py
- utils.py
- __init__.py
- llava.py
- llava_xtuner.py
- __init__.py
- model.py
- prompt.py
- __init__.py
- valley_eagle_chat.py
- __init__.py
- chat_uni_vi.py
- llama_vid.py
- pllava.py
- video_chatgpt.py
- video_llava.py
- videochat2.py
- __init__.py
- sharecaptioner.py
- xcomposer.py
- xcomposer2.py
- xcomposer2_4KHD.py
- xcomposer2d5.py
- __init__.py
- aria.py
- base.py
- bunnyllama3.py
- cambrian.py
- chameleon.py
- cogvlm.py
- deepseek_vl.py
- deepseek_vl2.py
- eagle_x.py
- emu.py
- falcon_vlm.py
- h2ovl_mississippi.py
- idefics.py
- instructblip.py
- internvl_chat.py
- janus.py
- kosmos.py
- llama_vision.py
- mantis.py
- mgm.py
- minicpm_v.py
- minigpt4.py
- minimonkey.py
- mixsense.py
- mmalaya.py
- molmo.py
- monkey.py
- moondream.py
- mplug_owl2.py
- mplug_owl3.py
- nvlm.py
- omchat.py
- omnilmm.py
- open_flamingo.py
- ovis.py
- paligemma.py
- pandagpt.py
- parrot.py
- phi3_vision.py
- pixtral.py
- points.py
- qh_360vl.py
- qwen_vl.py
- rbdash.py
- ross.py
- sail_vl.py
- slime.py
- smolvlm.py
- transcore_m.py
- vila.py
- vintern_chat.py
- visualglm.py
- vita.py
- vxverse.py
- wemm.py
- xgen_mm.py
- yi_vl.py
- __init__.py
- __main__.py
- config.py
- inference.py
- inference_mt.py
- inference_video.py
- run.py
- tools.py
- .pre-commit-config.yaml
- LICENSE
- Makefile
- MANIFEST.in
- README.md
- README_zh.md
- requirements.txt
- setup.cfg
- setup.py
- viz.py
- README.md
- data_pipeline.png
- example1.png
- exploration.png
- interleaving_reasoning.png
- IRG.png
- PTST.png
- reasoning_example.png
- reasoning_example1.png
- result_7B.png
- inference.sh
- vllm_deploy.sh
- vllm_inference.sh
- vllm_inference_local.sh
- fsdp_config.yaml
- ds_z0_config.json
- ds_z2_config.json
- ds_z2_offload_config.json
- ds_z3_config.json
- ds_z3_offload_config.json
- qwen2_full_sft.yaml
- llama3_full_sft.yaml
- llama3_full_sft.yaml
- llama3_lora_sft.yaml
- train.sh
- llama3_full_sft.yaml
- expand.sh
- llama3_freeze_sft.yaml
- llama3_lora_sft.yaml
- llama3_full_sft.yaml
- llama3_lora_predict.yaml
- init.sh
- llama3_lora_sft.yaml
- llama3.yaml
- llama3_full_sft.yaml
- llama3_lora_sft.yaml
- llama3_sglang.yaml
- llama3_vllm.yaml
- llava1_5.yaml
- qwen2_vl.yaml
- llama3_full_sft.yaml
- llama3_gptq.yaml
- llama3_lora_sft.yaml
- qwen2vl_lora_sft.yaml
- llama3_full_sft.yaml
- qwen2vl_full_sft.yaml
- llama3_lora_dpo.yaml
- llama3_lora_eval.yaml
- llama3_lora_kto.yaml
- llama3_lora_ppo.yaml
- llama3_lora_pretrain.yaml
- llama3_lora_reward.yaml
- llama3_lora_sft.yaml
- llama3_lora_sft_ds3.yaml
- llama3_lora_sft_ray.yaml
- llama3_preprocess.yaml
- llava1_5_lora_sft.yaml
- qwen2vl_lora_dpo.yaml
- qwen2vl_lora_sft.yaml
- llama3_lora_sft_aqlm.yaml
- llama3_lora_sft_awq.yaml
- llama3_lora_sft_bnb_npu.yaml
- llama3_lora_sft_gptq.yaml
- llama3_lora_sft_otfq.yaml
- README.md
- README_zh.md
- deepseed_node.sh
- train.yaml
- deepseed_node.sh
- train.yaml
- .gitignore
- inference.py
- README.md
- requirements.txt
- vllm_inference.py
- vllm_inference_local.py
# 설치 가이드
1. 코드 내려받기
git clone https://github.com/Osilly/Vision-R1
깃허브에서 프로젝트 코드 전체를 내 컴퓨터로 내려받습니다.
cd Vision-R1
방금 내려받은 프로젝트 폴더 안으로 이동합니다.
2. 공식 설치 스크립트
쉬움 추천사전 준비물
- Python 3 pip 명령어를 쓰려면 Python이 필요합니다.
pip install -U flash-attn --no-build-isolation
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
설치 후 새 터미널을 열고, 프로그램의 버전 확인 명령(예: --version)으로 정상 설치됐는지 확인하세요.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
3. Docker
쉬움사전 준비물
- Git GitHub에서 프로젝트 코드를 내려받으려면 필요합니다.
- Docker Desktop 컨테이너를 빌드하고 실행하려면 필요합니다. 설치 후 실행해서 백그라운드에 켜두세요.
⚠️ 이 프로젝트는 규모가 큰 저장소라, 이 방법이 실제 핵심 제품이 아니라 내부 하위 패키지를 가리키는 것일 수 있습니다. README 전체를 함께 확인해보세요.
docker build -f evaluation/evalscope/docker/dockerfile -t vision-r1 .
Dockerfile을 기반으로 실행 가능한 이미지를 빌드합니다.
docker run -p 8080:80 vision-r1
빌드된 이미지를 실제 컨테이너로 실행합니다.
터미널에 docker compose ps 를 입력해 컨테이너들이 Up 상태인지 확인하세요. README에 포트 번호가 적혀있다면 브라우저에서 http://localhost:포트번호 로 접속해보세요.
4. Python
쉬움사전 준비물
pip install -r requirements.txt
requirements.txt 등에 명시된 파이썬 라이브러리를 설치합니다.
pip install -U flash-attn --no-build-isolation
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
에러 메시지 없이 실행되고 터미널에 안내 문구가 출력되면 정상입니다.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
5. Make
보통사전 준비물
- Git GitHub에서 프로젝트 코드를 내려받으려면 필요합니다.
- Make Linux/macOS는 보통 기본 설치되어 있습니다. Windows는 별도 설치(예: MSYS2, WSL)가 필요합니다.
⚠️ 이 프로젝트는 규모가 큰 저장소라, 이 방법이 실제 핵심 제품이 아니라 내부 하위 패키지를 가리키는 것일 수 있습니다. README 전체를 함께 확인해보세요.
cd evaluation/evalscope
이 프로젝트의 관련 파일이 하위 폴더 안에 있어서, 먼저 그 폴더로 이동합니다.
make
생성된 빌드 설정을 바탕으로 실제 컴파일을 진행해 실행 파일을 만듭니다.
에러 없이 끝나면 성공입니다. 생성된 실행 파일을 직접 실행해보세요.
// repository documentation
Was this content helpful?
(0 ratings)
