https://www.mdu.se/

mdu.sePublications
4243444546474845 of 60
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
A Confidence-Aware Hybrid Vision-Language Framework for Food Recognition and Nutritional Monitoring
Kocaeli Univ, Dept Comp Engn, TR-41001 Izmit, Turkiye.
Kocaeli Univ, Dept Comp Engn, TR-41001 Izmit, Turkiye; Kocaeli Univ, Wireless Informat & Intelligent Syst WINS Res Ctr, TR-41001 Izmit, Turkiye.
Kocaeli Univ, Dept Comp Engn, TR-41001 Izmit, Turkiye; Kocaeli Univ, Wireless Informat & Intelligent Syst WINS Res Ctr, TR-41001 Izmit, Turkiye.
Kocaeli Univ, Wireless Informat & Intelligent Syst WINS Res Ctr, TR-41001 Izmit, Turkiye.
Show others and affiliations
2026 (English)In: Nutrients, E-ISSN 2072-6643, Vol. 18, no 15, article id 2449Article in journal (Refereed) Published
Abstract [en]

Background/Objectives: The objective evaluation of regional dietary intake remains a core challenge in personalized health management due to complex plate presentations and a lack of culturally specific dataset benchmarks. Methods: This study introduces a confidence-aware hybrid vision-language framework engineered for traditional Turkish food recognition and structured nutritional assessment. Results: We curate a balanced dataset containing 14,711 verified images spanning 40 representative Turkish culinary classes to train and evaluate seven deep learning architectures. Among the visual models, EfficientNet V2-L achieved the highest standalone performance with an accuracy of 93.47%, 0.92 macro-precision, 0.92 macro-recall, and a 0.92 F1 score. To overcome visual ambiguity and automate content analysis, a confidence-aware routing strategy escalates uncertain predictions ( tau<0.70 ) or user-rejected classifications to the Google Gemini 2.5 Flash multimodal large language model (MLLM). Conclusions: This hybrid paradigm yields a combined classification accuracy of 95.50% while validating portion weight estimations within a mean absolute error (MAE) of 18.42 g and total energy within 36.75 kcal. Fully realized as a cross-platform Flutter mobile application, the end-to-end pipeline demonstrates localized plate detection, adaptive portion analysis, and structured nutrient tracking, providing a scalable design for consumer-facing digital nutrition platforms.

Place, publisher, year, edition, pages
MDPI AG , 2026. Vol. 18, no 15, article id 2449
Keywords [en]
digital nutrition, food image recognition, nutrition estimation, multimodal large language models, confidence-aware routing, personalized dietary assessment, mobile health
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:mdh:diva-78828DOI: 10.3390/nu18152449ISI: 001848100100001PubMedID: 42588072Scopus ID: 2-s2.0-105047155589OAI: oai:DiVA.org:mdh-78828DiVA, id: diva2:2095501
Available from: 2026-08-26 Created: 2026-08-26 Last updated: 2026-08-26Bibliographically approved

Open Access in DiVA

fulltext(11739 kB)13 downloads
File information
File name FULLTEXT01.pdfFile size 11739 kBChecksum SHA-512
4ee977779b1a2d3b4ba2bcbf2eded15918b1cfd8fbb81a17cc8f44a450ec21fdc0919984786e1ffd494e407a83831c1c24e64362c1c01090130a3a4458f61c4d
Type fulltextMimetype application/pdf

Other links

Publisher's full textPubMedScopus

Authority records

Fotouhi, Hossein

Search in DiVA

By author/editor
Fotouhi, Hossein
By organisation
Department of Computer Science & Engineering
In the same journal
Nutrients
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

doi
pubmed
urn-nbn

Altmetric score

doi
pubmed
urn-nbn
Total: 389 hits
4243444546474845 of 60
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf