TY - GEN
T1 - A Multi-agent LLM System for Automated Requirements Analysis
T2 - Euromicro Conference on Software Engineering and Advanced Applications
AU - Sami, Malik Abdul
AU - Zhang, Zheying
AU - Waseem, Muhammad
AU - Kemell, Kai Kristian
AU - Rasheed, Zeeshan
AU - Herda, Tomas
AU - Hasan, Md Toufique
AU - Rasku, Jussi
AU - Abrahamsson, Pekka
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2025
Y1 - 2025
N2 - Manually specifying and prioritizing user stories in software projects is time-consuming and prone to inconsistency. This paper investigates whether Large Language Models (LLMs) can support these activities through a role-based multi-agent system. The proposed system uses four models: GPT-3.5 Turbo, GPT-4o, LLaMA 3.3, and Mistral-Nemo, to generate user stories from a project description and prioritize them using prompts simulating stakeholder roles. Relevance is evaluated using semantic similarity to the project description, and prioritization consistency is assessed using Kendall’s Tau distance against expert rankings. Results indicate that all models generate functionally relevant requirements with high semantic similarity, although clarity and conciseness vary. In prioritization, the models show moderate alignment with expert rankings, particularly for mid- and low-priority items, while variability across runs remains a challenge.
AB - Manually specifying and prioritizing user stories in software projects is time-consuming and prone to inconsistency. This paper investigates whether Large Language Models (LLMs) can support these activities through a role-based multi-agent system. The proposed system uses four models: GPT-3.5 Turbo, GPT-4o, LLaMA 3.3, and Mistral-Nemo, to generate user stories from a project description and prioritize them using prompts simulating stakeholder roles. Relevance is evaluated using semantic similarity to the project description, and prioritization consistency is assessed using Kendall’s Tau distance against expert rankings. Results indicate that all models generate functionally relevant requirements with high semantic similarity, although clarity and conciseness vary. In prioritization, the models show moderate alignment with expert rankings, particularly for mid- and low-priority items, while variability across runs remains a challenge.
KW - Large language model
KW - Multi-agent system
KW - Requirements generation
KW - Requirements prioritization
U2 - 10.1007/978-3-032-04200-2_12
DO - 10.1007/978-3-032-04200-2_12
M3 - Conference contribution
AN - SCOPUS:105016582791
SN - 9783032041999
T3 - Lecture Notes in Computer Science
SP - 178
EP - 187
BT - Software Engineering and Advanced Applications
A2 - Taibi, Davide
A2 - Smite, Darja
PB - Springer
Y2 - 10 September 2025 through 12 September 2025
ER -