Skip to main navigation Skip to search Skip to main content

Multimodal Generative AI Framework for Advancing Sustainable Development Goals

Research output: A Conference proceeding or a Chapter in BookConference contributionpeer-review

Abstract

The Sustainable Development Goals (SDGs) established by the United Nations provide an ambitious framework aimed at fostering a more just, sustainable, and prosperous global society. Recently, artificial intelligence (AI) has become a significant asset in striving towards these objectives. This paper outlines an innovative computational framework utilizing multimodal generative AI models to enhance assistive technologies that support various SDG initiatives. The proposed framework consists of three distinct architectural phases, each built upon novel model designs that merge image and text embeddings along with their combined representations. In Stage I, the architecture employs multimodal image-text embedding through contrastive learning and image pre-training, facilitating zero-shot classification. Stage II features a Multimodal RAG architecture that integrates CLIP with open-source large language models (LLMs), specifically leveraging GPT-2 and GPT-J in a cascading manner. Finally, Stage III utilizes a CLIP-assisted diffusion model that iteratively refines images through denoising random noise based on text prompts for high-quality image synthesis. Performance assessments for each stage were conducted using several open-source datasets and validated against image-text descriptions for specific use cases tied to the SDGs. The encouraging results from this study highlight the transformative capabilities of multimodal generative AI frameworks in furthering the United Nations’ Sustainable Development Goals by combining various data types—particularly text and images—for comprehensive insights and practical solutions.

Original languageEnglish
Title of host publicationNext-Generation Networks and Deployable Artificial Intelligence - Proceedings of NGNDAI 2025
EditorsDeepak Gupta, Dushyant Kumar Singh, Divya Kumar, Girija Chetty
PublisherSpringer
Pages309-323
Number of pages15
Volume2
ISBN (Print)9783032153944
DOIs
Publication statusPublished - 2026
EventInternational Conference on Next-Generation Networks and Deployable Artificial Intelligence, NGNDAI 2025 - Prayagraj, India
Duration: 18 Sept 202520 Sept 2025

Publication series

NameLecture Notes in Networks and Systems
Volume1793 LNNS
ISSN (Print)2367-3370
ISSN (Electronic)2367-3389

Conference

ConferenceInternational Conference on Next-Generation Networks and Deployable Artificial Intelligence, NGNDAI 2025
Country/TerritoryIndia
CityPrayagraj
Period18/09/2520/09/25

Fingerprint

Dive into the research topics of 'Multimodal Generative AI Framework for Advancing Sustainable Development Goals'. Together they form a unique fingerprint.

Cite this