https://www.mdu.se/

mdu.sePublications
1234562 of 6
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Data-Driven Analytics for Industrial Batch Processes: Integrating Batch Data Analytics, Machine Learning, Cost Sensitivity, and Post-hoc Analysis
Mälardalen University, Faculty of Engineering and Health Sciences, Department of Engineering Sciences.ORCID iD: 0000-0002-2455-3203
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

This thesis investigates how data-driven analytics can be systematically implemented in legacy industrial batch processes to support robust, interpretable, and operationally relevant decision-making.Industrial batch environments are characterised by heterogeneous data structures, variable process trajectories, evolving operating conditions, and limited contextualisation, complicating the direct application of conventional analytical methods.

The work develops an integrated analytical framework combining batch data analytics, machine learning, cost-sensitive learning, and model-agnostic post-hoc analysis.Batch data analytics is used to contextualise and consolidate irregular industrial process data into analytics-ready representations suitable for downstream modelling.Machine learning methods are subsequently applied to perform classification, regression, and degradation modelling across multiple industrial case studies, with emphasis placed on parsimonious and interpretable models rather than unnecessary model complexity.

To extend predictive modelling beyond conventional accuracy-oriented evaluation, cost-sensitive learning is introduced to align model behaviour with operational and economic objectives.The proposed framework demonstrates how predictive confidence and model coverage can be balanced against operational risk and cost constraints, enabling selective and value-aware decision-support strategies.

Beyond predictive optimisation, the thesis introduces a model-agnostic post-hoc analysis framework for examining model behaviour across operational regimes.By embedding interpretability metrics into reduced-dimensional representations and constructing continuous behavioural landscapes through surrogate modelling, regions associated with confidence, systematic error, uncertainty, and interaction-driven behaviour can be identified and analysed.The results demonstrate that predictive performance is not uniformly distributed across the input space, but instead governed by distinct operational regimes with varying levels of reliability and interpretability.

The framework is validated through multiple industrial case studies within alloy production, ceramic manufacturing, and degradation modelling of electrical resistance heating wires.The results show that structured batch contextualisation improves the suitability of industrial data for machine learning, that selective modelling strategies can achieve substantially higher predictive performance within identified operational regions, and that region-aware post-hoc analysis enables diagnostically grounded evaluation of model behaviour under changing industrial conditions.

The thesis contributes a coherent methodological framework for trustworthy industrial analytics in legacy batch environments by integrating structured data contextualisation, interpretable machine learning, value-aware evaluation, and region-aware behavioural analysis into a unified decision-support perspective.

Place, publisher, year, edition, pages
Västerås: Mälardalens universitet, 2026.
Series
Mälardalen University Press Dissertations, ISSN 1651-4238 ; 472
National Category
Energy Engineering
Research subject
Energy- and Environmental Engineering
Identifiers
URN: urn:nbn:se:mdh:diva-77868ISBN: 978-91-7485-762-7 (print)OAI: oai:DiVA.org:mdh-77868DiVA, id: diva2:2074497
Public defence
2026-08-20, Gamma, Mälardalens universitet, Västerås, 09:15 (English)
Opponent
Supervisors
Available from: 2026-06-18 Created: 2026-06-17 Last updated: 2026-06-18Bibliographically approved
List of papers
1. Comparing Feature and Trajectory-Based Remaining Useful Life Modeling of Electrical Resistance Heating Wires
Open this publication in new window or tab >>Comparing Feature and Trajectory-Based Remaining Useful Life Modeling of Electrical Resistance Heating Wires
2024 (English)In: Proc. Annu. Conf. Progn. Health Manag. Soc., PHM, PHM Society , 2024, no 1Conference paper, Published paper (Refereed)
Abstract [en]

Industrial heating significantly contributes to global greenhouse gas emissions, accounting for a substantial portion of annual emissions. The transition to fossil-free operations in the heating industry is closely linked to advancements in industrial electrical heating systems, especially those using resistance heating wires. In this context, Prognostics and Health Management is crucial for enhancing system reliability and sustainability through predictive maintenance strategies. The integration of machine learning technologies into Prognostics and Health Management has significantly improved the precision and applicability of Remaining Useful Life modeling. This improvement enables more accurate predictions of component lifespans, optimizes maintenance schedules, and enhances operational efficiency in industrial heating applications. These developments are essential for reducing greenhouse gas emissions in the sector. This paper serves as a guide for conducting Remaining Useful Life modeling for industrial batch processes. It evaluates and compares two methodologies: deep learning-based approaches using full time-series data, such as recurrent neural networks and their variants, and feature-engineering-based methods, including random forest regression and support vector machines. Our results show that the feature-oriented approach performs better overall in terms of predictive accuracy and computational efficiency. The study includes a detailed sensitivity analysis and hyperparameter estimation for each method, providing valuable insights into developing robust and transparent Prognostics and Health Management systems. These systems are crucial in supporting the heating industry’s move towards more sustainable and emission-free operations. The findings reveal that feature-oriented methods are both performant and robust, particularly excelling in handling outliers. The random forest regression model, in particular, demonstrated the highest performance on the test dataset according to the chosen evaluation metrics. Conversely, trajectory-oriented methods exhibited less bias across varying levels of degradation, a helpful characteristic for Prognostics and Health Management systems. While feature-oriented methods tend to systematically underestimate Remaining Useful Life at high true values and overestimate it at low actual values, this issue is less pronounced in trajectory-oriented models. Overall, these insights highlight the strengths and limitations of each approach, guiding the development of more effective and reliable predictive maintenance strategies.

Place, publisher, year, edition, pages
PHM Society, 2024
Series
Proceedings of the Annual Conference of the Prognostics and Health Management Society, PHM, ISSN 2325-0178
Keywords
Batch data processing, Diagnosis, Greenhouse gas emissions, Nuclear power plants, Recurrent neural networks, Support vector regression, Feature-oriented methods, Health management systems, Heating wire, Life models, Maintenance strategies, Predictive maintenance, Prognostic and health management, Random forests, Remaining useful lives
National Category
Civil Engineering
Identifiers
urn:nbn:se:mdh:diva-69256 (URN)10.36001/phmconf.2024.v16i1.3913 (DOI)2-s2.0-85210248759 (Scopus ID)9781936263059 (ISBN)
Conference
Proceedings of the Annual Conference of the Prognostics and Health Management Society, PHM
Available from: 2024-12-04 Created: 2024-12-04 Last updated: 2026-06-18Bibliographically approved
2. Trust, but Verify-Post-Hoc Analysis of Industrial Machine Learning via Interpretability Metric Embedding and Surrogate Mapping
Open this publication in new window or tab >>Trust, but Verify-Post-Hoc Analysis of Industrial Machine Learning via Interpretability Metric Embedding and Surrogate Mapping
2026 (English)In: Sensors, E-ISSN 1424-8220, Vol. 26, no 10, article id 3232Article in journal (Refereed) Published
Abstract [en]

In industrial machine learning, predictive performance alone is insufficient to ensure reliable deployment, as model behaviour may vary across different regions of the input space under limited data and evolving process conditions. This work investigates whether such variation can be systematically analysed through post-hoc methods. A model-agnostic framework is proposed in which interpretability metrics, including residuals and feature attributions, are embedded into a low-dimensional space and approximated using a continuous surrogate model. This representation enables the analysis of model behaviour as a structured landscape, rather than as isolated pointwise explanations. The approach is applied to ceramic heating element production, where two distinct regimes are identified. One corresponds to a stable region with consistent and accurate predictions, while the other reflects a transitional regime associated with increased ambiguity and sensitivity to feature interactions. These regimes are shown to align with known process conditions and temporal variation. The results demonstrate that model behaviour can be organised into coherent regions that are not observable through aggregate performance metrics alone. This provides a structured basis for post-hoc analysis, supporting targeted interpretation and further investigation of model reliability in industrial settings.

Place, publisher, year, edition, pages
MDPI AG, 2026
Keywords
post-hoc analysis, explainable AI, UMAP, interpretability metrics, industrial machine learning, decision landscape
National Category
Computer Sciences
Identifiers
urn:nbn:se:mdh:diva-77162 (URN)10.3390/s26103232 (DOI)001775516600001 ()42198039 (PubMedID)2-s2.0-105040114136 (Scopus ID)
Available from: 2026-06-03 Created: 2026-06-03 Last updated: 2026-06-18Bibliographically approved
3. Consolidating industrial batch process data for machine learning
Open this publication in new window or tab >>Consolidating industrial batch process data for machine learning
2022 (English)In: / [ed] Esko Juuso, Bernt Lie, Erik Dahlquist and Jari Ruuska, Linkoping University Electronic Press , 2022, p. 76-83Conference paper, Published paper (Refereed)
Abstract [en]

The paradigm change of Industry 4.0 brings attention to data-driven modeling and the incentive to apply machine learning methods in the process industry. Further, capitalizing on a great deal of data available is an adverse task. For batch processes, the dataset is in a threeway format (Batch × Sensor × Time). Depending on the process and the goal of the analysis, it might be necessary to aggregate batches together. For this reason, a campaign unfolding structure is applied. By grouping the batches under new labels relevant to the analytical goal, campaigns are created. These labels can be created from periodical occurrences, such as refurbishing the refractory lining in the case of the case study. In order to utilize the three-way batch format, it is necessary to align the batches. In order to address this, the feature-oriented approach Statistical Pattern Analysis (SPA) is applied. SPA derives statistics, e.g., mean, skewness and kurtosis from the time series, consequently aligning the batches. The SPA and the campaign approach create a dataset consisting of select statistics instead of an irregular three-way array. Functional data analysis (FDA) is used to smooth and extract first- and second-order derivative information from the sensors in which functional behavior can be observed before creating features. Principal Component Analysis (PCA) is used to examine the final dataset. Further, industrial processes are notoriously nonlinear, and even more so batch processes. Therefore, kernel-based principal component analysis (KPCA) is used to review the final dataset. The KPCA can accommodate different underlying characteristics by modifying the kernel function used. 

Place, publisher, year, edition, pages
Linkoping University Electronic Press, 2022
Series
Linköping Electronic Conference Proceedings, ISSN 1650-3686, E-ISSN 1650-3740 ; 185
Keywords
Batch Process Analysis (BDA), Batch preprocessing, Functional Data Analysis (FDA), Statistical Pattern Analysis (SPA), Kernel Principal Component Analysis (KPCA)
National Category
Computer Systems
Identifiers
urn:nbn:se:mdh:diva-61115 (URN)10.3384/ecp21185 (DOI)978-91-7929-219-5 (ISBN)
Conference
The First SIMS EUROSIM Conference on Modelling and Simulation, SIMS EUROSIM 2021, and 62nd International Conference of Scandinavian Simulation Society, SIMS 2021, September 21-23, Virtual Conference, Finland
Available from: 2022-12-06 Created: 2022-12-06 Last updated: 2026-06-18Bibliographically approved
4. Cost-Sensitive Decision Support for Industrial Batch Processes
Open this publication in new window or tab >>Cost-Sensitive Decision Support for Industrial Batch Processes
2023 (English)In: Sensors, E-ISSN 1424-8220, Vol. 23, no 23, article id 9464Article in journal (Refereed) Published
Abstract [en]

In this work, cost-sensitive decision support was developed. Using Batch Data Analytics (BDA) methods of the batch data structure and feature accommodation, the batch process property and sensor data can be accommodated. The batch data structure organises the batch processes' data, and the feature accommodation approach derives statistics from the time series, consequently aligning the time series with the other features. Three machine learning classifiers were implemented for comparison: Logistic Regression (LR), Random Forest Classifier (RFC), and Support Vector Machine (SVM). It is possible to filter out the low-probability predictions by leveraging the classifiers' probability estimations. Consequently, the decision support has a trade-off between accuracy and coverage. Cost-sensitive learning was used to implement a cost matrix, which further aggregates the accuracy-coverage trade into cost metrics. Also, two scenarios were implemented for accommodating out-of-coverage batches. The batch is discarded in one scenario, and the other is processed. The Random Forest classifier was shown to outperform the other classifiers and, compared to the baseline scenario, had a relative cost of 26%. This synergy of methods provides cost-aware decision support for analysing the intricate workings of a multiprocess batch data system.

Place, publisher, year, edition, pages
MDPI, 2023
Keywords
Batch Data Analytics (BDA), feature-oriented, cost-sensitive learning, decision support, machine learning
National Category
Energy Engineering
Identifiers
urn:nbn:se:mdh:diva-65129 (URN)10.3390/s23239464 (DOI)001115971100001 ()38067837 (PubMedID)2-s2.0-85179140167 (Scopus ID)
Available from: 2023-12-20 Created: 2023-12-20 Last updated: 2026-06-18Bibliographically approved
5. Evaluating Modelling Performance: Sensitivity Analysis of Data Volume in Industrial Batch Processes
Open this publication in new window or tab >>Evaluating Modelling Performance: Sensitivity Analysis of Data Volume in Industrial Batch Processes
2024 (English)In: Proceedings of the Second SIMS EUROSIM Conference on Modelling and Simulation, SIMS EUROSIM 2024 / [ed] Esko Juuso, Jari Ruuska, Gaurav Mirlekar and Lars Eriksson, Linkoping University Electronic Press , 2024, p. 415-424Conference paper, Published paper (Refereed)
Abstract [en]

This study conducts a sensitivity analysis to evaluate the influence of varying datavolumes on model performance within multi-product batch processes in the iron and steelindustry. Nine machine learning models, encompassing both ensemble and parametric methods,were rigorously tested using a data withholding approach. The results demonstrate thatensemble models, particularly Random Forest and Gradient Boosting, consistently outperformedparametric models across different data volumes, showcasing superior generalisation androbustness to outliers. These findings underscore the importance of careful model selection andcomprehensive data preprocessing in enhancing model performance and suggest that ensemblemethods are particularly well-suited for complex industrial applications where data quality andvolume are critical.

Place, publisher, year, edition, pages
Linkoping University Electronic Press, 2024
Series
Linköping Electronic Conference Proceedings, ISSN 1650-3686, E-ISSN 1650-3740 ; 212
Keywords
Machine Learning, Model selection, Performance evaluation, Data volumesensitivity, Iron and steel industry, Industrial batch processes
National Category
Computer Vision and Learning Systems
Identifiers
urn:nbn:se:mdh:diva-77866 (URN)10.3384/ecp212.057 (DOI)978-91-8075-984-7 (ISBN)
Conference
Second SIMS EUROSIM Conference on Modelling and Simulation, SIMS EUROSIM 2024, Oulu, Finland, 11-12 September, 2024
Funder
Knowledge Foundation
Available from: 2026-06-17 Created: 2026-06-17 Last updated: 2026-06-29Bibliographically approved
6. Deriving degradation drivers in electrical heating wires through Model-agnostic Post-hoc analysis
Open this publication in new window or tab >>Deriving degradation drivers in electrical heating wires through Model-agnostic Post-hoc analysis
2026 (English)In: Proceedings of the 15th european conference on industrial furnaces and boilers, 2026Conference paper, Published paper (Refereed)
Abstract [en]

This paper applies a post-hoc analytical framework to investigate degradation behaviour in industrialelectrical heating systems. A feature-based representation of resistance trajectories is used to train aremaining useful life (rul) model, after which interpretability metrics are mapped onto a reduced latentspace to examine model confidence, systematic deviation and lifetime-dependent structure. The analysisreveals distinct operational regimes corresponding to early, middle and late degradation phases, andhighlights how local and global statistical descriptors contribute differently to the separation of thesephases. The results demonstrate how post-hoc model analysis can support the extraction of processrelevant mechanisms and provide structured guidance for model interpretation in industrial settings.

National Category
Computer Vision and Learning Systems
Identifiers
urn:nbn:se:mdh:diva-77864 (URN)zenodo.org/records/20595383 (DOI)
Conference
INFUB-15, Porto, Portugal, 7–10 April 2026
Funder
Knowledge Foundation
Available from: 2026-06-17 Created: 2026-06-17 Last updated: 2026-06-18Bibliographically approved

Open Access in DiVA

No full text in DiVA

Authority records

Mählkvist, Simon

Search in DiVA

By author/editor
Mählkvist, Simon
By organisation
Department of Engineering Sciences
Energy Engineering

Search outside of DiVA

GoogleGoogle Scholar

isbn
urn-nbn

Altmetric score

isbn
urn-nbn
Total: 33 hits
1234562 of 6
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf