Vision-Language Assistive Navigation for Visually Impaired Using BLIP Fine-Tuning
Charan Sai Ponnada, Dr. S. Kumar, R. Patel
International Symposium on Advanced Electrical and Communication Technologies (ISAECT) · November 1, 2025
BLEU Score Gain
+18%
Dataset Size
50K pairs
LoRA Stages
3-Stage
Inference Speed
2.5x
Abstract
This paper presents a novel approach to assistive navigation for visually impaired individuals using vision-language models. We fine-tune the BLIP (Bootstrapping Language-Image Pre-training) model using a three-stage LoRA (Low-Rank Adaptation) strategy to generate contextual navigation descriptions from visual input. Our method achieves significant improvements in BLEU scores and inference speed compared to baseline approaches, demonstrating the effectiveness of parameter-efficient fine-tuning for assistive AI applications.
Keywords
Citation
@inproceedings{ponnada2025vision,
title={Vision-Language Assistive Navigation for Visually Impaired Using BLIP Fine-Tuning},
author={Ponnada, Charan Sai and Kumar, S. and Patel, R.},
booktitle={Proceedings of ISAECT 2025},
year={2025}
}