https://www.mdu.se/

mdu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Comparative analysis of text mining and clustering techniques for assessing functional dependency between manual test cases
Mälardalen University, School of Innovation, Design and Engineering, Innovation and Product Realisation. Ericsson AB, Compute Platforms Engn Unit, Stockholm, Sweden.ORCID iD: 0000-0002-8724-9049
Mälardalen University, School of Innovation, Design and Engineering, Innovation and Product Realisation.ORCID iD: 0000-0003-0073-1674
German Aerosp Ctr, Inst Software Technol, Cologne, Germany; Univ Cologne, Dept Math & Comp Sci, Cologne, Germany.
Show others and affiliations
2025 (English)In: Software quality journal, ISSN 0963-9314, E-ISSN 1573-1367, Vol. 33, no 2, article id 24Article in journal (Refereed) Published
Abstract [en]

Text mining techniques, particularly those leveraging machine learning for natural language processing, have gained significant attention for qualitative data analysis in software testing. However, their complexity and lack of transparency can pose challenges, especially in safety-critical domains where simpler, interpretable solutions are often preferred unless accuracy is heavily compromised. This study investigates the trade-offs between complexity, effort, accuracy, and utility in text mining and clustering techniques, focusing on their application for detecting functional dependencies among manual integration test cases in safety-critical systems. Using empirical data from an industrial testing project at ALSTOM Sweden, we evaluate various string distance methods, NCD compressors, and machine learning approaches. The results highlight the impact of preprocessing techniques, such as tokenization, and intrinsic factors, such as text length, on algorithm performance. Findings demonstrate how text mining and clustering can be optimized for safety-critical contexts, offering actionable insights for researchers and practitioners aiming to balance simplicity and effectiveness in their testing workflows.

Place, publisher, year, edition, pages
Springer Nature , 2025. Vol. 33, no 2, article id 24
Keywords [en]
Artificial intelligence, Clustering, Natural language processing, Text mining, Software testing
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:mdh:diva-71444DOI: 10.1007/s11219-025-09722-7ISI: 001489598700001Scopus ID: 2-s2.0-105005412458OAI: oai:DiVA.org:mdh-71444DiVA, id: diva2:1960725
Available from: 2025-05-23 Created: 2025-05-23 Last updated: 2026-03-30Bibliographically approved

Open Access in DiVA

fulltext(1396 kB)2 downloads
File information
File name FULLTEXT01.pdfFile size 1396 kBChecksum SHA-512
98fbb164bf6e08c81943d14fafede9eeb4d484b04e3bb66999cc197a96b410871e01a7b73d4ecaca9703a663dbc0253a714c2256b70098a3a96dead3728da35c
Type fulltextMimetype application/pdf

Other links

Publisher's full textScopus

Authority records

Tahvili, SaharHatvani, LeoAfzal, Wasif

Search in DiVA

By author/editor
Tahvili, SaharHatvani, LeoAfzal, Wasif
By organisation
Innovation and Product RealisationEmbedded Systems
In the same journal
Software quality journal
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 2 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 175 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf