https://www.mdu.se/

mdu.sePublications
4546474849505148 of 60
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
LATENT OPTIMIZATION DYNAMICS FOR DETECTION OF AI-GENERATED IMAGES
Mälardalen University.
2026 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

AI models based on latent diffusion models (LDMs), such as Stable Diffusion, FLUX, and Midjourney, have made it possible for virtually anyone to generate photorealistic images with arbitrary content. This rapid development has created a growing need for reliable methods to detect AI-generated images.

Many popular AI models rely on a pretrained variational autoencoder (VAE) to reduce computational costs and improve performance. Several existing detection methods exploit the fact that AI-generated images can be reconstructed more accurately through a VAE than real photographs. Since the models that these detection methods were trained to identify were primarily optimized to generate aesthetically pleasing images, it is possible that part of their detection capability relies on an inherent aesthetic bias rather than on more general traces of AI generation. Over time, even humans have learned to recognize the "perfect" AI aesthetic that has characterized earlier models.

Newer models, such as Flux, Qwen, and Nano Banana, have increasingly moved away from this idealized aesthetic and instead focus on generating images with a high degree of realism. For example, they can produce images that appear to have been taken with a shaky smartphone, exhibiting motion blur and various visual imperfections. These models now also can be prompted to edit existing images by altering only parts of the original image, which increases the potential for their misuse in deepfakes.

This development has placed greater demands on AI detectors, and methods that relied solely on the reconstruction error from a single forward pass through the VAE may struggle to detect images that lack a distinct AI aesthetic.

This thesis investigates whether optimizing the latent representation through gradient descent affects reconstruction error differently depending on whether an image is a real photograph or AI-generated. By studying the change in reconstruction error over repeated optimization steps using image gradients as loss function, detection  seem to improve over previous methods on newer AI-generative models. 

Place, publisher, year, edition, pages
2026.
National Category
Artificial Intelligence
Identifiers
URN: urn:nbn:se:mdh:diva-78742OAI: oai:DiVA.org:mdh-78742DiVA, id: diva2:2093137
Subject / course
Computer Science
Supervisors
Examiners
Available from: 2026-08-25 Created: 2026-08-18 Last updated: 2026-08-25Bibliographically approved

Open Access in DiVA

fulltext(69316 kB)20 downloads
File information
File name FULLTEXT01.pdfFile size 69316 kBChecksum SHA-512
f62fc8751ceee9f73b8aed989cd9609c077b03f864390fc2ebfa5f15eec635b3f1843a26161b7e53cd722688970ee4eb1a4eabeffa5227975e41f1d710a4e5e2
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Dahlgren, Martin
By organisation
Mälardalen University
Artificial Intelligence

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 690 hits
4546474849505148 of 60
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf