Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation
Official PyTorch implementation for our paper Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation (AAAI 2026).
uv venv --python 3.10
uv sync --dev
uv pip install -e .Continue here to setup the synthetic pretraining dataset with incomplete task-language coverage.
Continue here to setup the finetuning dataset based on Multi30K, COCO Karpathy, DOCCI, ImageParagraphs.
We need to preprocess the model checkpoints.
- Remap the decoder embeddings for Florence-2 to match the Gemma-2 tokenizer.
- Insert the Florence-2 vision backbone and cross-attention layers into Gemma-2. These steps allow us to load the model faster.
python src/prepare.pyThe script generates tokenizer mappings and an init_state_dict.pt file for the Florence-2 base and large models, as well as the Gemma-2-2B and Gemma-2-9B models with the Florence-2 large vision backbone.
The folder slurm/ contains training scripts for the different model sizes for continous pretraining and finetuning.
To fit the scaling laws, we measure cross-entropy loss on a synthetic evaluation set.
python src/eval_loss.py
--batch_size 8 # Batch size for eval
--eval_dataset_root_path {} # Eval dataset path data/...
--dataset_stats_file {} # The dataset stats file data/tok_counts.json
--pretrained_path {} # The checkpoint path
--training_steps 10_000 # The number of training steps 500, 2000, 5000, 10000
--training_batch_size 1024 # Batch size used for pre-training
--model_name {} # Choose microsoft/Florence-2-base, microsoft/Florence-2-large, florence-gemma-2-large, florence-gemma-2-xlargeRun evaluation on Multi30K, COCO Karpathy, DOCCI, ImageParagraphs, CoMMuTE, and XM3600.
python src/eval.py
--model_name {} # Name of the model
--model_type {} # Model type
--input_res 768 # 224 or 768
--model_path {} # Path to fine-tuned model
--batch_size 8 # Batch size for eval
--eval_task_set full # Task set selection
--use_prefix # Append a prefix to the decoder inputOur implementation is based on transformers, Gemma 2, and Florence 2. This research has been funded by the Federal Ministry of Research, Technology and Space of Germany under grant no. 01IS23004B RIDMI and 01IS22094C WEST-AI. Computational resources were provided by the German AI Service Center WestAI.
@inproceedings{spravil2026scaling,
title={Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation},
author={Spravil, Julian and Houben, Sebastian and Behnke, Sven},
booktitle={Proceedings of the 40th AAAI Conference on Artificial Intelligence},
year={2026}
}
