Article

Conversational Image Recognition Chatbot Using Vision-Language Models

Author : 1Dr. P. Harini, 2M. Triveni, 3M. Hema Srivalli, 4N. Sudheer Vali

This project presents a Conversational Image Recognition Chatbot that integrates Artificial Intelligence, Computer Vision, and Vision-Language Models to provide intelligent image analysis through natural language conversations. The system enables users to upload an image and interact with it by asking questions about its contents, such as identifying objects, describing scenes, extracting text, and understanding contextual relationships. It ensures accurate and real-time image understanding by leveraging advanced deep learning techniques and multimodal AI models. Unlike traditional image recognition systems that only classify objects or generate static captions, the proposed system supports interactive, context-aware conversations, allowing users to ask follow-up questions and receive detailed responses. Conventional image recognition applications often lack conversational capabilities, contextual reasoning, and personalised interaction, making them less effective for complex visual understanding tasks. There is a need for an intelligent, user-friendly, and cost-effective image recognition system that combines computer vision with conversational artificial intelligence to deliver accurate, explainable, and human-like responses. The proposed chatbot addresses these limitations by providing an efficient and scalable solution suitable for applications in education, healthcare, tourism, accessibility, customer support, and smart digital assistants.


Full Text Attachment
//