Testing of Pre-trained Large Language Models for Zero-shot Software Engineering Tasks
Implementing Organization
Indian Institute Of Technology, Gandhinagar
Principal Investigator
Dr. Shouvick Mondal
Indian Institute Of Technology, Gandhinagar
shouvick.mondal@iitgn.ac.in
Project Overview
Pre-trained Large Language Models (LLMs) such as ChatGPT, Gemini, and Llama are extensively utilised to automate tasks at our workplace, and the pace at which tasks get completed is highly appreciable. However, their quality needs to be critically assessed before being utilised in a safe and trustworthy manner. The recent surge in the application of prompt-driven LLMs for Software Engineering (SE) tasks (e.g. code generation, code testing, code summarization) brings out two major concerns: (i) incompleteness of software testing (problem #1) in the era of LLMs, (ii) testing of genAI systems, i.e., LLMs themselves (problem #2). In this project, we focus on the black box testing of LLMs by analysing their responses on zero-shot prompts. Zero-shot learning (prompting) is a Machine Learning technique that enables models to generalise and make predictions on classes or tasks that were not seen during the training phase. This amounts to testing how well the LLMs perform SE tasks when the instruction from the developer does not contain any example of a valid/safe response that it must produce. Existing LLM based approaches are capable of handling only tasks for specific phases of the Software Development Life Cycle (SDLC) instead of offering an end-to-end support. We propose to fulfil this research gap by seeking answers to the following high-level research question (RQ): "LLM4SE with zero-shots: how far across different SDLC phases can we go?". By testing LLMs along these lines, we will uncover adversarial (ill formed) patterns in LLM-prompt manipulations that can potentially prevent LLMs from producing their intended response in terms of deliverables for a particular phase of the SDLC. Of all the phases, the software (code) testing phase alone consumes more than 50% of the maintenance cost and over 80% of the end-to-end time of the overall SDLC. In this regard, we will focus on several functional/non-functional aspects of software testing by and for LLMs, and identify bottlenecks. Currently, LLM-driven SDLC processes focus only on specific phases of the SDLC with none providing end-to-end support spanning all its phases. We would like to investigate how the application of zero-shot prompts on a set of few LLMs can achieve full coverage of all SDLC phases, to the extent that fully autonomous SE is possible in the near future. Successful implementation of this project will have several outcomes: (i) identification/ detection/ classification of LLM bugs and feedback to LLM developers for continuous improvement, (ii) development of novel prompt engineering strategies tailored to specific or a group of SDLC activities, (iii) contribution to Government of India initiatives like Digital India, IndiaAI, and Free and Open Source Software (FOSS). A strong use case would be to integrate our developed technology with an open collaborative development platform like OpenForge for e-governance applications.
Disclaimer:
Information available on this portal is sourced from various organizations and is provided for informational purposes only. Users are advised to verify details from the respective official sources.
Please enter your details
Please provide your name and email to continue. Your details are saved in this browser for future use.
Latest Updates
Loading…
⚠️
You are leaving this website
You are about to be redirected to an external website that is not operated by
India Science, Technology & Innovation (ISTI) Portal.