Hello, during the VQA testing, I used the second-stage pre-training weights and conducted one round of training on the LLAVA OneVision dataset for evaluation. I found that some metrics differed significantly from the results reported in the paper (such as OCRBench, DocVQA, etc.). Could you please provide detailed information about your training configuration, or make the corresponding evaluation weights publicly available?
Hello, during the VQA testing, I used the second-stage pre-training weights and conducted one round of training on the LLAVA OneVision dataset for evaluation. I found that some metrics differed significantly from the results reported in the paper (such as OCRBench, DocVQA, etc.). Could you please provide detailed information about your training configuration, or make the corresponding evaluation weights publicly available?