ocr-vqgan
OCR-VQGAN, a discrete image encoder (tokenizer and detokenizer) for figure images in Paper2Fig100k dataset. Implementation of OCR Perceptual loss for clear text-within-image generation. Fork from VQGAN in CompVis/taming-transformers
파일 탐색기
최종 버전 다운로드 (.zip)- ocr_v2.png
- original.png
- reconstruction.png
- table_comparison_1.PNG
- table_comparison_2.PNG
- ocr-vqgan-f16-c16384-d256.yaml
- ocr-vqgan-imagenet-16384.yaml
- compute_ocr_perceptual_loss.py
- compute_ssim.py
- evaluate_DALLE_VQVAE.py
- generate_qualitative_results.py
- parse_ICDAR2013_img_to_VQGAN.py
- parse_paper2fig1_img_to_VQGAN.py
- plot_ocr_features.py
- prepare_eval_samples.py
- base.py
- custom.py
- image_transforms.py
- utils.py
- vqgan.py
- model.py
- model.py
- __init__.py
- craft.py
- lpips.py
- vgg16_bn.py
- vqperceptual.py
- quantize.py
- util.py
- lr_scheduler.py
- util.py
- .gitignore
- environment.yaml
- main.py
- README.md
- setup.py
// repository documentation
Was this content helpful?
(0 ratings)
