Multi-Modality-Arena
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
파일 탐색기
최종 버전 다운로드 (.zip)파일 수가 많아 일부만 표시됩니다. 전체 파일은 위 다운로드 버튼으로 확인해 주세요.
- tiny_lvlm_ehub_6_12.png
- figure1.png
- figure2.png
- figure3.png
- figure4.png
- all.css
- cropper.css
- bad.png
- better.png
- Tie.png
- 1.png
- 2.png
- 3.png
- 4.png
- 5.png
- A_better.png
- B_better.png
- bad.png
- clear.png
- Tie.png
- A_better.png
- B_better.png
- bad.png
- clear.png
- Tie.png
- 1.jpg
- 2.jpg
- 3.jpg
- api.svg
- Built.svg
- crop.svg
- demo.jpg
- demo1.jpg
- full.svg
- Git.svg
- head_full.svg
- link.svg
- logo.svg
- lvlm-ehub.png
- messge.svg
- move.svg
- Opengvlab_LOGO.svg
- pause.svg
- photo.svg
- photo_white.svg
- player.svg
- Rotate.svg
- Scale.svg
- Send.svg
- upload.svg
- video.svg
- video_white.svg
- Wechat.jpeg
- cropper.js
- gpt.js
- jquery-3.1.1.min.js
- header.html
- index.html
- kun_basketball.jpg
- merlion.png
- tiananmen.jpg
- __init__.py
- decoders.py
- extended.py
- image_net.py
- image_net_22k.py
- adapters.py
- distributed.py
- loaders.py
- logging.py
- metrics.py
- samplers.py
- utils.py
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- defaults.yaml
- defaults_dial.yaml
- defaults_cap.yaml
- defaults_ret.yaml
- defaults_vqa.yaml
- eval_vqa.yaml
- defaults_12m.yaml
- defaults_3m.yaml
- defaults_ret.yaml
- defaults.yaml
- balanced_testdev.yaml
- balanced_val.yaml
- defaults.yaml
- defaults.yaml
- defaults_2B_multi.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults_ret.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults_cap.yaml
- defaults_caption.yaml
- defaults_vqa.yaml
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_instruct_flant5xl.yaml
- blip2_instruct_flant5xxl.yaml
- blip2_instruct_vicuna13b.yaml
- blip2_instruct_vicuna7b.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- img2prompt_vqa_base.yaml
- pnp_vqa_3b.yaml
- pnp_vqa_base.yaml
- pnp_vqa_large.yaml
- unifiedqav2_3b_config.json
- unifiedqav2_base_config.json
- unifiedqav2_large_config.json
- albef_classification_ve.yaml
- albef_feature_extractor.yaml
- albef_nlvr.yaml
- albef_pretrain_base.yaml
- albef_retrieval_coco.yaml
- albef_retrieval_flickr.yaml
- albef_vqav2.yaml
- alpro_qa_msrvtt.yaml
- alpro_qa_msvd.yaml
- alpro_retrieval_didemo.yaml
- alpro_retrieval_msrvtt.yaml
- bert_config.json
- bert_config_alpro.json
- blip_caption_base_coco.yaml
- blip_caption_large_coco.yaml
- blip_classification_base.yaml
- blip_feature_extractor_base.yaml
- blip_itm_base.yaml
- blip_itm_large.yaml
- blip_nlvr.yaml
- blip_pretrain_base.yaml
- blip_pretrain_large.yaml
- blip_retrieval_coco.yaml
- blip_retrieval_flickr.yaml
- blip_vqa_aokvqa.yaml
- blip_vqa_okvqa.yaml
- blip_vqav2.yaml
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- gpt_dialogue_base.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_t5_instruct.py
- blip2_vicuna_instruct.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip.py
- blip_caption.py
- blip_classification.py
- blip_feature_extractor.py
- blip_image_text_matching.py
- blip_nlvr.py
- blip_outputs.py
- blip_pretrain.py
- blip_retrieval.py
- blip_vqa.py
- nlvr_encoder.py
- __init__.py
- conv2d_same.py
- features.py
- helpers.py
- linear.py
- vit.py
- vit_utils.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- __init__.py
- __init__.py
- llama.py
- llama_adapter.py
- tokenizer.py
- utils.py
- adapt_tokenizer.py
- attention.py
- blocks.py
- configuration_mpt.py
- hf_prefixlm_converter.py
- meta_init_context.py
- modeling_mpt.py
- norm.py
- param_init_fns.py
- __init__.py
- apply_delta.py
- consolidate.py
- llava.py
- llava_mpt.py
- make_delta.py
- utils.py
- __init__.py
- constants.py
- conversation.py
- utils.py
- __init__.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- minigpt4.yaml
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- base_model.py
- blip2.py
- blip2_outputs.py
- eva_vit.py
- mini_gpt4.py
- modeling_llama.py
- Qformer.py
- __init__.py
- base_processor.py
- blip_processors.py
- randaugment.py
- __init__.py
- minigpt4_eval.yaml
- __init__.py
- configuration_mplug_owl.py
- modeling_mplug_owl.py
- processing_mplug_owl.py
- tokenization_mplug_owl.py
- __init__.py
- config.json
- configuration_otter.py
- flamingo_pt2otter_hf.py
- modeling_otter.py
- otter_pt2otter_hf.py
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5base_vitL.yaml
- blip2_pretrain_flant5large_vitL.yaml
- blip2_pretrain_flant5small_vitL.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt1.3b_vitL.yaml
- blip2_pretrain_opt125m_vitL.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt2.7b_vitL.yaml
- blip2_pretrain_opt350m_vitL.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_opt6.7b_vitL.yaml
- blip2_pretrain_vicuna7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- bert_config.json
- bert_config_alpro.json
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_llama.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_vicuna.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip_outputs.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- transforms_video.py
- __init__.py
- vpgtrans_demo.yaml
- __init__.py
- test_blip2.py
- test_instructblip.py
- test_llama_adapter_v2.py
- test_llava.py
- test_minigpt4.py
- test_mplug_owl.py
- test_otter.py
- test_vpgtrans.py
- __init__.py
- call_gpt.py
- vcr_conv.py
- vcr_prompt.py
- ve_conv.py
- ve_prompt.py
- __init__.py
- vcr.py
- ve.py
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- defaults.yaml
- defaults_dial.yaml
- defaults_cap.yaml
- defaults_ret.yaml
- defaults_vqa.yaml
- eval_vqa.yaml
- defaults_12m.yaml
- defaults_3m.yaml
- defaults_ret.yaml
- defaults.yaml
- balanced_testdev.yaml
- balanced_val.yaml
- defaults.yaml
- defaults.yaml
- defaults_2B_multi.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults_ret.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults_cap.yaml
- defaults_caption.yaml
- defaults_vqa.yaml
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_instruct_flant5xl.yaml
- blip2_instruct_flant5xxl.yaml
- blip2_instruct_vicuna13b.yaml
- blip2_instruct_vicuna7b.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- img2prompt_vqa_base.yaml
- pnp_vqa_3b.yaml
- pnp_vqa_base.yaml
- pnp_vqa_large.yaml
- unifiedqav2_3b_config.json
- unifiedqav2_base_config.json
- unifiedqav2_large_config.json
- albef_classification_ve.yaml
- albef_feature_extractor.yaml
- albef_nlvr.yaml
- albef_pretrain_base.yaml
- albef_retrieval_coco.yaml
- albef_retrieval_flickr.yaml
- albef_vqav2.yaml
- alpro_qa_msrvtt.yaml
- alpro_qa_msvd.yaml
- alpro_retrieval_didemo.yaml
- alpro_retrieval_msrvtt.yaml
- bert_config.json
- bert_config_alpro.json
- blip_caption_base_coco.yaml
- blip_caption_large_coco.yaml
- blip_classification_base.yaml
- blip_feature_extractor_base.yaml
- blip_itm_base.yaml
- blip_itm_large.yaml
- blip_nlvr.yaml
- blip_pretrain_base.yaml
- blip_pretrain_large.yaml
- blip_retrieval_coco.yaml
- blip_retrieval_flickr.yaml
- blip_vqa_aokvqa.yaml
- blip_vqa_okvqa.yaml
- blip_vqav2.yaml
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- gpt_dialogue_base.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_t5_instruct.py
- blip2_vicuna_instruct.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip.py
- blip_caption.py
- blip_classification.py
- blip_feature_extractor.py
- blip_image_text_matching.py
- blip_nlvr.py
- blip_outputs.py
- blip_pretrain.py
- blip_retrieval.py
- blip_vqa.py
- nlvr_encoder.py
- __init__.py
- conv2d_same.py
- features.py
- helpers.py
- linear.py
- vit.py
- vit_utils.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- __init__.py
- __init__.py
- blip2_lib.py
- instructblip_lib.py
- llama.py
- llama_adapter.py
- LLama_lib.py
- llava_lib.py
- minigpt4_config.py
- minigpt4_lib.py
- model_mae.py
- mplugowl_lib.py
- otter_lib.py
- tokenizer.py
- utils.py
- vgtrans_lib.py
- checklist.chk
- params.json
- README.md
- tokenizer.model
- models_mae.py
- models_mae_.py
- bpe_simple_vocab_16e6.txt.gz
- __init__.py
- helpers.py
- imagebind_model.py
- multimodal_preprocessors.py
- transformer.py
- CODE_OF_CONDUCT.md
- CONTRIBUTING.md
- data.py
- demo.py
- LICENSE
- model_card.md
- README.md
- requirements.txt
- __init__.py
- llama.py
- llama_adapter.py
- tokenizer.py
- utils.py
- __init__.py
- adapt_tokenizer.py
- attention.py
- blocks.py
- configuration_mpt.py
- hf_prefixlm_converter.py
- meta_init_context.py
- modeling_mpt.py
- norm.py
- param_init_fns.py
- __init__.py
- apply_delta.py
- consolidate.py
- llava.py
- llava_mpt.py
- make_delta.py
- utils.py
- __init__.py
- constants.py
- conversation.py
- utils.py
- __init__.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- minigpt4.yaml
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- base_model.py
- blip2.py
- blip2_outputs.py
- eva_vit.py
- mini_gpt4.py
- modeling_llama.py
- Qformer.py
- __init__.py
- base_processor.py
- blip_processors.py
- randaugment.py
- __init__.py
- minigpt4_eval.yaml
- vcr_val_random500_annoid.yaml
- ve_dev_random500_pairid.yaml
- __init__.py
- blip_caption.py
- check_sample.py
- sample_data.py
- sample_data_ve.py
- __init__.py
- configuration_mplug_owl.py
- modeling_mplug_owl.py
- processing_mplug_owl.py
- tokenization_mplug_owl.py
- __init__.py
- config.json
- configuration_otter.py
- flamingo_pt2otter_hf.py
- modeling_otter.py
- otter_pt2otter_hf.py
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5base_vitL.yaml
- blip2_pretrain_flant5large_vitL.yaml
- blip2_pretrain_flant5small_vitL.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt1.3b_vitL.yaml
- blip2_pretrain_opt125m_vitL.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt2.7b_vitL.yaml
- blip2_pretrain_opt350m_vitL.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_opt6.7b_vitL.yaml
- blip2_pretrain_vicuna7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- bert_config.json
- bert_config_alpro.json
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_llama.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_vicuna.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip_outputs.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- transforms_video.py
- __init__.py
- vpgtrans_demo.yaml
- blip_gpt_main.py
- environment.yml
- README.md
- vcr_eval.py
- ve_eval.py
- blip2_opt.py
- color.csv
- component.csv
- material.csv
- others.csv
- shape.csv
- ImageNet_mapping.txt
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- defaults.yaml
- defaults_dial.yaml
- defaults_cap.yaml
- defaults_ret.yaml
- defaults_vqa.yaml
- eval_vqa.yaml
- defaults_12m.yaml
- defaults_3m.yaml
- defaults_ret.yaml
- defaults.yaml
- balanced_testdev.yaml
- balanced_val.yaml
- defaults.yaml
- defaults.yaml
- defaults_2B_multi.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults_ret.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults_cap.yaml
- defaults_caption.yaml
- defaults_vqa.yaml
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_instruct_flant5xl.yaml
- blip2_instruct_flant5xxl.yaml
- blip2_instruct_vicuna13b.yaml
- blip2_instruct_vicuna7b.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- img2prompt_vqa_base.yaml
- pnp_vqa_3b.yaml
- pnp_vqa_base.yaml
- pnp_vqa_large.yaml
- unifiedqav2_3b_config.json
- unifiedqav2_base_config.json
- unifiedqav2_large_config.json
- albef_classification_ve.yaml
- albef_feature_extractor.yaml
- albef_nlvr.yaml
- albef_pretrain_base.yaml
- albef_retrieval_coco.yaml
- albef_retrieval_flickr.yaml
- albef_vqav2.yaml
- alpro_qa_msrvtt.yaml
- alpro_qa_msvd.yaml
- alpro_retrieval_didemo.yaml
- alpro_retrieval_msrvtt.yaml
- bert_config.json
- bert_config_alpro.json
- blip_caption_base_coco.yaml
- blip_caption_large_coco.yaml
- blip_classification_base.yaml
- blip_feature_extractor_base.yaml
- blip_itm_base.yaml
- blip_itm_large.yaml
- blip_nlvr.yaml
- blip_pretrain_base.yaml
- blip_pretrain_large.yaml
- blip_retrieval_coco.yaml
- blip_retrieval_flickr.yaml
- blip_vqa_aokvqa.yaml
- blip_vqa_okvqa.yaml
- blip_vqav2.yaml
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- gpt_dialogue_base.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_t5_instruct.py
- blip2_vicuna_instruct.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip.py
- blip_caption.py
- blip_classification.py
- blip_feature_extractor.py
- blip_image_text_matching.py
- blip_nlvr.py
- blip_outputs.py
- blip_pretrain.py
- blip_retrieval.py
- blip_vqa.py
- nlvr_encoder.py
- __init__.py
- conv2d_same.py
- features.py
- helpers.py
- linear.py
- vit.py
- vit_utils.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- __init__.py
- checklist.chk
- params.json
- README.md
- tokenizer.model
- models_mae.py
- adapt_tokenizer.py
- attention.py
- blocks.py
- configuration_mpt.py
- hf_prefixlm_converter.py
- meta_init_context.py
- modeling_mpt.py
- norm.py
- param_init_fns.py
- __init__.py
- apply_delta.py
- consolidate.py
- llava.py
- llava_mpt.py
- make_delta.py
- utils.py
- __init__.py
- constants.py
- conversation.py
- utils.py
- __init__.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- minigpt4.yaml
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- base_model.py
- blip2.py
- blip2_outputs.py
- eva_vit.py
- mini_gpt4.py
- modeling_llama.py
- Qformer.py
- __init__.py
- base_processor.py
- blip_processors.py
- randaugment.py
- __init__.py
- minigpt4_eval.yaml
- __init__.py
- configuration_mplug_owl.py
- modeling_mplug_owl.py
- processing_mplug_owl.py
- tokenization_mplug_owl.py
- __init__.py
- config.json
- configuration_otter.py
- flamingo_pt2otter_hf.py
- modeling_otter.py
- otter_pt2otter_hf.py
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5base_vitL.yaml
- blip2_pretrain_flant5large_vitL.yaml
- blip2_pretrain_flant5small_vitL.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt1.3b_vitL.yaml
- blip2_pretrain_opt125m_vitL.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt2.7b_vitL.yaml
- blip2_pretrain_opt350m_vitL.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_opt6.7b_vitL.yaml
- blip2_pretrain_vicuna7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- bert_config.json
- bert_config_alpro.json
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_llama.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_vicuna.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip_outputs.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- transforms_video.py
- __init__.py
- vpgtrans_demo.yaml
- ImageNetVC_blip2.py
- ImageNetVC_instructblip.py
- ImageNetVC_llama.py
- ImageNetVC_llava.py
- ImageNetVC_minigpt4.py
- ImageNetVC_mplugowl.py
- ImageNetVC_otter.py
- ImageNetVC_vpgtrans.py
- Readme.md
- dist_eval_knn.sh
- eval_model.sh
- __init__.py
- caption_datasets.py
- cls_datasets.py
- embod_datasets.py
- formula_datasets.py
- kie_datasets.py
- ocr_datasets.py
- vqa_datasets.py
- whoops.py
- __init__.py
- caption.py
- cider.py
- classification.py
- embodied.py
- kie.py
- mrr.py
- ocr.py
- tools.py
- vqa.py
- LVLM Evaluation on Embodied Benchmark.xlsx
- README.md
- coco_mci.json
- coco_oc.json
- vcr1_mci.json
- vcr1_oc.json
- coco_pope_adversarial.json
- coco_pope_adversarial1.json
- coco_pope_popular.json
- coco_pope_popular1.json
- coco_pope_random.json
- coco_pope_random1.json
- meta.py
- eval.py
- eval_knn.py
- owl_eval.py
- README.md
- requirements.txt
- whoops_vqa_bem_eval.py
- blip2_opt.py
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- defaults.yaml
- defaults_dial.yaml
- defaults_cap.yaml
- defaults_ret.yaml
- defaults_vqa.yaml
- eval_vqa.yaml
- defaults_12m.yaml
- defaults_3m.yaml
- defaults_ret.yaml
- defaults.yaml
- balanced_testdev.yaml
- balanced_val.yaml
- defaults.yaml
- defaults.yaml
- defaults_2B_multi.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults_ret.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults_cap.yaml
- defaults_caption.yaml
- defaults_vqa.yaml
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_instruct_flant5xl.yaml
- blip2_instruct_flant5xxl.yaml
- blip2_instruct_vicuna13b.yaml
- blip2_instruct_vicuna7b.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- img2prompt_vqa_base.yaml
- pnp_vqa_3b.yaml
- pnp_vqa_base.yaml
- pnp_vqa_large.yaml
- unifiedqav2_3b_config.json
- unifiedqav2_base_config.json
- unifiedqav2_large_config.json
- albef_classification_ve.yaml
- albef_feature_extractor.yaml
- albef_nlvr.yaml
- albef_pretrain_base.yaml
- albef_retrieval_coco.yaml
- albef_retrieval_flickr.yaml
- albef_vqav2.yaml
- alpro_qa_msrvtt.yaml
- alpro_qa_msvd.yaml
- alpro_retrieval_didemo.yaml
- alpro_retrieval_msrvtt.yaml
- bert_config.json
- bert_config_alpro.json
- blip_caption_base_coco.yaml
- blip_caption_large_coco.yaml
- blip_classification_base.yaml
- blip_feature_extractor_base.yaml
- blip_itm_base.yaml
- blip_itm_large.yaml
- blip_nlvr.yaml
- blip_pretrain_base.yaml
- blip_pretrain_large.yaml
- blip_retrieval_coco.yaml
- blip_retrieval_flickr.yaml
- blip_vqa_aokvqa.yaml
- blip_vqa_okvqa.yaml
- blip_vqav2.yaml
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- gpt_dialogue_base.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_t5_instruct.py
- blip2_vicuna_instruct.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip.py
- blip_caption.py
- blip_classification.py
- blip_feature_extractor.py
- blip_image_text_matching.py
- blip_nlvr.py
- blip_outputs.py
- blip_pretrain.py
- blip_retrieval.py
- blip_vqa.py
- nlvr_encoder.py
- __init__.py
- conv2d_same.py
- features.py
- helpers.py
- linear.py
- vit.py
- vit_utils.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- __init__.py
- checklist.chk
- params.json
- README.md
- tokenizer.model
- models_mae.py
- adapt_tokenizer.py
- attention.py
- blocks.py
- configuration_mpt.py
- hf_prefixlm_converter.py
- meta_init_context.py
- modeling_mpt.py
- norm.py
- param_init_fns.py
- __init__.py
- apply_delta.py
- consolidate.py
- llava.py
- llava_mpt.py
- make_delta.py
- utils.py
- __init__.py
- constants.py
- conversation.py
- utils.py
- model_med_eval_sp.py
- __init__.py
- apply_delta.py
- consolidate.py
- llava.py
- make_delta.py
- utils.py
- __init__.py
- constants.py
- conversation.py
- openai_api.py
- utils.py
- synpic57813.jpg
- dataset.py
- test.py
- test.sh
- __init__.py
- utils.py
- install.sh
- requirements.txt
- dataset.py
- randaugment.py
- blocks.py
- vqa_model.py
- blocks.py
- pmc_clip.py
- timm_model.py
- utils.py
- __init__.py
- adalora.py
- adaption_prompt.py
- lora.py
- p_tuning.py
- prefix_tuning.py
- prompt_tuning.py
- __init__.py
- adapters_utils.py
- config.py
- other.py
- save_and_load.py
- __init__.py
- import_utils.py
- mapping.py
- peft_model.py
- test.py
- test.sh
- LICENSE
- PMC-VQA
- __init__.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- minigpt4.yaml
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- base_model.py
- blip2.py
- blip2_outputs.py
- eva_vit.py
- mini_gpt4.py
- modeling_llama.py
- Qformer.py
- __init__.py
- base_processor.py
- blip_processors.py
- randaugment.py
- __init__.py
- minigpt4_eval.yaml
- __init__.py
- configuration_mplug_owl.py
- modeling_mplug_owl.py
- processing_mplug_owl.py
- tokenization_mplug_owl.py
- __init__.py
- config.json
- configuration_otter.py
- flamingo_pt2otter_hf.py
- modeling_otter.py
- otter_pt2otter_hf.py
- config.json
- special_tokens_map.json
- tokenizer.model
- tokenizer_config.json
- config.json
- special_tokens_map.json
- tokenizer.json
- tokenizer_config.json
- vocab.txt
- __init__.py
- blocks.py
- helpers.py
- multimodality_model.py
- my_embedding_layer.py
- position_encoding.py
- transformer_decoder.py
- utils.py
- vit_3d.py
- test.py
- test.sh
- view1_frontal.jpg
- __init__.py
- binary.py
- caption_prompt.json
- case_report.py
- chestxray.py
- cls_prompt.json
- mammo_prompt.json
- MedPix_dataset.py
- modality_prompt.json
- paper_inline.py
- pmcoa.py
- pmcvqa.py
- radiology_feature_prompt.json
- radiopaedia.py
- report_prompt.json
- spinexr_prompt.json
- yes_no_prompt.json
- multi_dataset.py
- multi_dataset_test.py
- multi_dataset_test_for_close.py
- __init__.py
- blocks.py
- helpers.py
- multimodality_model.py
- my_embedding_layer.py
- position_encoding.py
- transformer_decoder.py
- utils.py
- vit_3d.py
- trainer.py
- datasampler.py
- test.py
- train.py
- config.json
- special_tokens_map.json
- tokenizer.model
- tokenizer_config.json
- config.json
- special_tokens_map.json
- tokenizer.json
- tokenizer_config.json
- vocab.txt
- __init__.py
- blocks.py
- helpers.py
- multimodality_model.py
- my_embedding_layer.py
- position_encoding.py
- transformer_decoder.py
- utils.py
- vit_3d.py
- dataset.py
- test.py
- test.sh
- requirements.txt
- run_eval_loss.sh
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5base_vitL.yaml
- blip2_pretrain_flant5large_vitL.yaml
- blip2_pretrain_flant5small_vitL.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt1.3b_vitL.yaml
- blip2_pretrain_opt125m_vitL.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt2.7b_vitL.yaml
- blip2_pretrain_opt350m_vitL.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_opt6.7b_vitL.yaml
- blip2_pretrain_vicuna7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- bert_config.json
- bert_config_alpro.json
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_llama.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_vicuna.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip_outputs.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- transforms_video.py
- __init__.py
- vpgtrans_demo.yaml
- medical_blip2.py
- medical_instructblip.py
- medical_llama_adapter2.py
- medical_llava.py
- medical_minigpt4.py
- medical_otter.py
- medical_owl.py
- medical_vpgtrans.py
- model_med_eval.py
- __init__.py
- apply_delta.py
- consolidate.py
- llava.py
- make_delta.py
- utils.py
- __init__.py
- constants.py
- conversation.py
- utils.py
- synpic57813.jpg
- dataset.py
- test.py
- test.sh
- __init__.py
- utils.py
- install.sh
- requirements.txt
- dataset.py
- randaugment.py
- blocks.py
- QA_model.py
- QA_model_mlp.py
- transformer.py
- test.py
- test.sh
- LICENSE
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- defaults.yaml
- vqg.yaml
- defaults_cap.yaml
- defaults_vqa.yaml
- eval_vqa.yaml
- vqg.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- pretrain_558.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- vqg.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- blip2_instruct_vicuna7b.yaml
- bliva_flant5xxl.yaml
- bliva_vicuna7b.yaml
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- base_model.py
- blip2.py
- blip2_vicuna_instruct.py
- bliva_flant5xxl.py
- bliva_vicuna7b.py
- bliva_vicuna7b_lora.py
- clip_vit.py
- eva_vit.py
- modeling_llama.py
- modeling_t5.py
- pretrain_bliva_flant5.py
- pretrain_bliva_vicuna7b.py
- Qformer.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- randaugment.py
- __init__.py
- runner_base.py
- runner_iter.py
- __init__.py
- base_task.py
- image_text_pretrain.py
- __init__.py
- __init__.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- cheetah_llama2.yaml
- cheetah_vicuna.yaml
- default.yaml
- __init__.py
- conversation.py
- conversation_llama2.py
- __init__.py
- base_model.py
- blip2.py
- blip2_outputs.py
- cheetah_llama2.py
- cheetah_vicuna.py
- eva_vit.py
- modeling_llama.py
- modeling_llama2.py
- Qformer.py
- __init__.py
- base_processor.py
- blip_processors.py
- randaugment.py
- __init__.py
- cheetah_eval_llama2.yaml
- cheetah_eval_vicuna.yaml
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- defaults.yaml
- defaults_dial.yaml
- defaults_cap.yaml
- defaults_ret.yaml
- defaults_vqa.yaml
- eval_vqa.yaml
- defaults_12m.yaml
- defaults_3m.yaml
- defaults_ret.yaml
- defaults.yaml
- balanced_testdev.yaml
- balanced_val.yaml
- defaults.yaml
- defaults.yaml
- defaults_2B_multi.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults_ret.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults_cap.yaml
- defaults_caption.yaml
- defaults_vqa.yaml
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_instruct_flant5xl.yaml
- blip2_instruct_flant5xxl.yaml
- blip2_instruct_vicuna13b.yaml
- blip2_instruct_vicuna7b.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- img2prompt_vqa_base.yaml
- pnp_vqa_3b.yaml
- pnp_vqa_base.yaml
- pnp_vqa_large.yaml
- unifiedqav2_3b_config.json
- unifiedqav2_base_config.json
- unifiedqav2_large_config.json
- albef_classification_ve.yaml
- albef_feature_extractor.yaml
- albef_nlvr.yaml
- albef_pretrain_base.yaml
- albef_retrieval_coco.yaml
- albef_retrieval_flickr.yaml
- albef_vqav2.yaml
- alpro_qa_msrvtt.yaml
- alpro_qa_msvd.yaml
- alpro_retrieval_didemo.yaml
- alpro_retrieval_msrvtt.yaml
- bert_config.json
- bert_config_alpro.json
- blip_caption_base_coco.yaml
- blip_caption_large_coco.yaml
- blip_classification_base.yaml
- blip_feature_extractor_base.yaml
- blip_itm_base.yaml
- blip_itm_large.yaml
- blip_nlvr.yaml
- blip_pretrain_base.yaml
- blip_pretrain_large.yaml
- blip_retrieval_coco.yaml
- blip_retrieval_flickr.yaml
- blip_vqa_aokvqa.yaml
- blip_vqa_okvqa.yaml
- blip_vqav2.yaml
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- gpt_dialogue_base.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- blip2.py
- blip2_t5_instruct.py
- blip2_vicuna_instruct.py
- modeling_llama.py
- modeling_t5.py
- Qformer.py
- __init__.py
- conv2d_same.py
- features.py
- helpers.py
- linear.py
- vit.py
- vit_utils.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- __init__.py
- __init__.py
- bpe_simple_vocab_16e6.txt.gz
- clip.py
- model.py
- simple_tokenizer.py
- apply_delta.py
- base_prompt.py
- datasets.py
- lars.py
- lr_decay.py
- lr_sched.py
- misc.py
- pos_embed.py
- quantization.py
- __init__.py
- eval_model.py
- generator.py
- mm_adaptation.py
- mm_adapter.py
- model.py
- tokenizer.py
- __init__.py
- llama.py
- llama_adapter.py
- tokenizer.py
- utils.py
- adapt_tokenizer.py
- attention.py
- blocks.py
- configuration_mpt.py
- hf_prefixlm_converter.py
- meta_init_context.py
- modeling_mpt.py
- norm.py
- param_init_fns.py
- __init__.py
- apply_delta.py
- consolidate.py
- llava.py
- llava_mpt.py
- make_delta.py
- utils.py
- __init__.py
- constants.py
- conversation.py
- utils.py
- LYNX.yaml
- Open_VQA_images.jsonl
- Open_VQA_videos.jsonl
- __init__.py
- __init__.py
- eval_datasets.py
- utils.py
- ablation.png
- logo_plus.png
- lynx.png
- open_vqa_image_result.png
- result_other.png
- LICENSE-adapter-transformers.txt
- LICENSE-BAAI-VISION.txt
- LICENSE-flamingo-pytorch.txt
- LICENSE-transformers.txt
- __init__.py
- resampler.py
- __init__.py
- configuration_llama.py
- convert_llama_weights_to_hf.py
- modeling_llama.py
- tokenization_llama.py
- __init__.py
- __init__.py
- eva_vit.py
- __init__.py
- adapter_modeling.py
- lynx.py
- generate.py
- generate.sh
- utils.py
- .DS_Store
- __init__.py
- configuration_blip_2.py
- convert_blip_2_original_to_pytorch.py
- modeling_blip_2.py
- processing_blip_2.py
- __init__.py
- configuration_instructblip.py
- convert_instructblip_original_to_pytorch.py
- modeling_instructblip.py
- processing_instructblip.py
- .DS_Store
- utils.py
- __init__.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- minigpt4.yaml
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- base_model.py
- blip2.py
- blip2_outputs.py
- eva_vit.py
- mini_gpt4.py
- modeling_llama.py
- Qformer.py
- __init__.py
- base_processor.py
- blip_processors.py
- randaugment.py
- __init__.py
- minigpt4_eval.yaml
- __init__.py
- configuration_mplug_owl.py
- modeling_mplug_owl.py
- processing_mplug_owl.py
- tokenization_mplug_owl.py
- __init__.py
- config.json
- configuration_otter.py
- flamingo_pt2otter_hf.py
- modeling_otter.py
- otter_pt2otter_hf.py
- __init__.py
- config.json
- configuration_otter.py
- converting_otter_fp32_to_fp16.py
- converting_otter_pt_to_hf.py
- flamingo_pt2otter_hf.py
- modeling_otter.py
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5base_vitL.yaml
- blip2_pretrain_flant5large_vitL.yaml
- blip2_pretrain_flant5small_vitL.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt1.3b_vitL.yaml
- blip2_pretrain_opt125m_vitL.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt2.7b_vitL.yaml
- blip2_pretrain_opt350m_vitL.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_opt6.7b_vitL.yaml
- blip2_pretrain_vicuna7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- bert_config.json
- bert_config_alpro.json
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_llama.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_vicuna.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip_outputs.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- transforms_video.py
- __init__.py
- vpgtrans_demo.yaml
- __init__.py
- test_automodel.py
- test_blip2.py
- test_bliva.py
- test_cheetah.py
- test_instructblip.py
- test_lavin.py
- test_llama_adapter_v2.py
- test_llava.py
- test_lynx.py
- test_mic.py
- test_minigpt4.py
- test_mplug_owl.py
- test_OFv2.py
- test_otter.py
- test_otter_image.py
- test_pandagpt.py
- test_vpgtrans.py
- config.json
- special_tokens_map.json
- tokenizer.model
- tokenizer_config.json
- config.json
- special_tokens_map.json
- tokenizer.json
- tokenizer_config.json
- vocab.txt
- __init__.py
- blocks.py
- helpers.py
- multimodality_model.py
- my_embedding_layer.py
- position_encoding.py
- transformer_decoder.py
- utils.py
- vit_3d.py
- test.py
- test.sh
- view1_frontal.jpg
- __init__.py
- binary.py
- caption_prompt.json
- case_report.py
- chestxray.py
- cls_prompt.json
- mammo_prompt.json
- MedPix_dataset.py
- modality_prompt.json
- paper_inline.py
- pmcoa.py
- pmcvqa.py
- radiology_feature_prompt.json
- radiopaedia.py
- report_prompt.json
- spinexr_prompt.json
- yes_no_prompt.json
- multi_dataset.py
- multi_dataset_test.py
- multi_dataset_test_for_close.py
- __init__.py
- blocks.py
- helpers.py
- multimodality_model.py
- my_embedding_layer.py
- position_encoding.py
- transformer_decoder.py
- utils.py
- vit_3d.py
- trainer.py
- datasampler.py
- test.py
- train.py
- config.json
- special_tokens_map.json
- tokenizer.model
- tokenizer_config.json
- config.json
- special_tokens_map.json
- tokenizer.json
- tokenizer_config.json
- vocab.txt
- __init__.py
- blocks.py
- helpers.py
- multimodality_model.py
- my_embedding_layer.py
- position_encoding.py
- transformer_decoder.py
- utils.py
- vit_3d.py
- dataset.py
- test.py
- test.sh
- requirements.txt
- run_eval.sh
- medical_datasets.py
- medicalqa.py
- eval_medical.py
- requirements.txt
- README.md
- test_path.json
- README.md
- model.jpg
- __init__.py
- config.json
- configuration_flamingo.py
- converting_flamingo_to_pytorch.py
- modeling_flamingo.py
- __init__.py
- vqa.py
- vqa_eval.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- defaults.yaml
- defaults_dial.yaml
- defaults_cap.yaml
- defaults_ret.yaml
- defaults_vqa.yaml
- eval_vqa.yaml
- defaults_12m.yaml
- defaults_3m.yaml
- defaults_ret.yaml
- defaults.yaml
- balanced_testdev.yaml
- balanced_val.yaml
- defaults.yaml
- defaults.yaml
- defaults_2B_multi.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults_ret.yaml
- defaults_cap.yaml
- defaults_qa.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults.yaml
- defaults_cap.yaml
- defaults_caption.yaml
- defaults_vqa.yaml
- blip2_caption_flant5xl.yaml
- blip2_caption_opt2.7b.yaml
- blip2_caption_opt6.7b.yaml
- blip2_coco.yaml
- blip2_instruct_flant5xl.yaml
- blip2_instruct_flant5xxl.yaml
- blip2_instruct_vicuna13b.yaml
- blip2_instruct_vicuna7b.yaml
- blip2_pretrain.yaml
- blip2_pretrain_flant5xl.yaml
- blip2_pretrain_flant5xl_vitL.yaml
- blip2_pretrain_flant5xxl.yaml
- blip2_pretrain_llama7b.yaml
- blip2_pretrain_opt2.7b.yaml
- blip2_pretrain_opt6.7b.yaml
- blip2_pretrain_vitL.yaml
- RN101-quickgelu.json
- RN101.json
- RN50-quickgelu.json
- RN50.json
- RN50x16.json
- RN50x4.json
- timm-efficientnetv2_rw_s.json
- timm-resnet50d.json
- timm-resnetaa50d.json
- timm-resnetblur50.json
- timm-swin_base_patch4_window7_224.json
- timm-vit_base_patch16_224.json
- timm-vit_base_patch32_224.json
- timm-vit_small_patch16_224.json
- ViT-B-16-plus-240.json
- ViT-B-16-plus.json
- ViT-B-16.json
- ViT-B-32-plus-256.json
- ViT-B-32-quickgelu.json
- ViT-B-32.json
- ViT-g-14.json
- ViT-H-14.json
- ViT-H-16.json
- ViT-L-14-280.json
- ViT-L-14-336.json
- ViT-L-14.json
- ViT-L-16-320.json
- ViT-L-16.json
- img2prompt_vqa_base.yaml
- pnp_vqa_3b.yaml
- pnp_vqa_base.yaml
- pnp_vqa_large.yaml
- unifiedqav2_3b_config.json
- unifiedqav2_base_config.json
- unifiedqav2_large_config.json
- albef_classification_ve.yaml
- albef_feature_extractor.yaml
- albef_nlvr.yaml
- albef_pretrain_base.yaml
- albef_retrieval_coco.yaml
- albef_retrieval_flickr.yaml
- albef_vqav2.yaml
- alpro_qa_msrvtt.yaml
- alpro_qa_msvd.yaml
- alpro_retrieval_didemo.yaml
- alpro_retrieval_msrvtt.yaml
- bert_config.json
- bert_config_alpro.json
- blip_caption_base_coco.yaml
- blip_caption_large_coco.yaml
- blip_classification_base.yaml
- blip_feature_extractor_base.yaml
- blip_itm_base.yaml
- blip_itm_large.yaml
- blip_nlvr.yaml
- blip_pretrain_base.yaml
- blip_pretrain_large.yaml
- blip_retrieval_coco.yaml
- blip_retrieval_flickr.yaml
- blip_vqa_aokvqa.yaml
- blip_vqa_okvqa.yaml
- blip_vqav2.yaml
- clip_resnet50.yaml
- clip_vit_base16.yaml
- clip_vit_base32.yaml
- clip_vit_large14.yaml
- clip_vit_large14_336.yaml
- gpt_dialogue_base.yaml
- med_config.json
- med_config_albef.json
- med_large_config.json
- default.yaml
- __init__.py
- blip2.py
- blip2_image_text_matching.py
- blip2_opt.py
- blip2_qformer.py
- blip2_t5.py
- blip2_t5_instruct.py
- blip2_vicuna_instruct.py
- modeling_llama.py
- modeling_opt.py
- modeling_t5.py
- Qformer.py
- __init__.py
- blip.py
- blip_caption.py
- blip_classification.py
- blip_feature_extractor.py
- blip_image_text_matching.py
- blip_nlvr.py
- blip_outputs.py
- blip_pretrain.py
- blip_retrieval.py
- blip_vqa.py
- nlvr_encoder.py
- __init__.py
- conv2d_same.py
- features.py
- helpers.py
- linear.py
- vit.py
- vit_utils.py
- __init__.py
- base_model.py
- clip_vit.py
- eva_vit.py
- med.py
- vit.py
- __init__.py
- base_processor.py
- blip_processors.py
- clip_processors.py
- functional_video.py
- gpt_processors.py
- randaugment.py
- __init__.py
- models_mae.cpython-38.pyc
- models_mae.py
- adapt_tokenizer.py
- attention.py
- blocks.py
- configuration_mpt.py
- hf_prefixlm_converter.py
- meta_init_context.py
- modeling_mpt.py
- norm.py
- param_init_fns.py
- __init__.py
- apply_delta.py
- consolidate.py
- llava.py
- llava_mpt.py
- make_delta.py
- utils.py
- __init__.py
- constants.py
- conversation.py
- utils.py
- __init__.py
- config.py
- dist_utils.py
- gradcam.py
- logger.py
- optims.py
- registry.py
- utils.py
- minigpt4.yaml
- default.yaml
- __init__.py
- conversation.py
- __init__.py
- base_model.py
- Qformer.py
- __init__.py
- minigpt4_eval.yaml
- __init__.py
- .gitignore
- controller.py
- demo.py
- env.yaml
- model_worker.py
- README.md
# 설치 가이드
1. 코드 내려받기
git clone https://github.com/OpenGVLab/Multi-Modality-Arena
깃허브에서 프로젝트 코드 전체를 내 컴퓨터로 내려받습니다.
cd Multi-Modality-Arena
방금 내려받은 프로젝트 폴더 안으로 이동합니다.
2. 공식 설치 스크립트
쉬움 추천사전 준비물
- Python 3 pip 명령어를 쓰려면 Python이 필요합니다.
pip install numpy gradio uvicorn fastapi
PyPI에 배포된 패키지를 바로 설치합니다. 소스 클론이 필요 없습니다.
설치 후 새 터미널을 열고, 프로그램의 버전 확인 명령(예: --version)으로 정상 설치됐는지 확인하세요.
이 레포의 README에 적힌 실제 명령어를 그대로 가져왔습니다.
// repository documentation
Was this content helpful?
(0 ratings)
