Across Domains

Generative AI "Across Domains"

Generative AI models are used to assist professionals from various fields in performing different tasks.

Select one of the areas below that is closest to your field of expertise and explore various tasks that can be supported by generative AI models:

Medicine
 
 

Can generative AI models produce more training data for other AI models in healthcare?

 

The more data there is, the better the deep learning model can become (Simon et al. 2024). However, medical data is often too scarce in this world to train modern deep neural networks. The reasons for this are the high costs of image acquisition and processing, strict data protection laws, and the low incidence of some diseases (Kazerouni et al. 2023). If generative AI models can be used to generate similar images based on the limited available training data, we may be able to improve all downstream AI models that can be used for diagnosis.

Generative adversarial models (GANs) first attracted attention by enriching medical training data with high-quality AI-generated images, thus laying the foundation for further revolutionary strategies for a variety of pathological image analysis tasks (Barragán-Montero et al. 2023; Suganyadevi et al. 2022). However, such augmentation techniques using GANs faced unstable training due to limited quality and diversity (Kazerouni et al. 2023).

Diffusion models have proven to be more efficient than GANs in image generation (Dhariwal and Nichol 2021). These models have been used to generate images to combat data scarcity (Müller-Franzes et al. 2021; Pinaya et al. 2022; Kazerouni et al. 2023). Models have also been developed that not only generate synthetic data from scarce real data but also accompany text inputs to improve quality. Kidder integrated the Stable Diffusion framework, an image synthesis module with latent diffusion models (Rombach et al. 2021), with DreamBooth, a text-to-image diffusion model (Ruiz et al. 2022), to create synthetic medical MRI and X-ray images (Kidder, 2024).

 

Left:
A. Synthetic MRI images of meningioma tumors generated using the Kidder model.
B. Synthetic MRI images of glioma tumors generated using the same model.

Right:
A. Actual training data from the CDD-CESM database (Kahled et al., 2022).
B. Synthetic mammography images generated by Kidder’s model based on text-to-image synthesis.

It has been shown that Kidder’s model generates synthetic medical images much faster than GANs, with GANs requiring 24–26 hours to train, while Kidder’s modified DreamBooth model only needed 10–15 minutes. Such text-input diffusion models can not only train diagnostic models that otherwise would not be trainable but also reduce the economic costs of collecting training data and the ethical concerns associated with using sensitive personal data for model training.

 
 
 
 
 
 
REFERENCES
  1. Simon, J. B. (2024). More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory. ICLR 2024 Poster. https://doi.org/https://openreview.net/forum?id=OdpIjS0vkO
  2. Kazerouni A, Aghdam EK, Heidari M et al.  Diffusion models in medical imaging: a comprehensive survey. Med Image Anal 2023;88:102846.
  3. Barragán-Montero A, Javaid U, Valdés G et al.  Artificial intelligence and machine learning for medical imaging: a technology review. Phys Med 2021;83:242–56.
  4. Suganyadevi S, Seethalakshmi V, Balasamy K. A review on deep learning in medical image analysis. Int J Multimed Inf Retr 2022;11:19–38.
  5. Dhariwal P, Nichol A. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems 2021;34:8780–94.
  6. Müller-Franzes G, Niehues J, Khader F et al.  Diffusion probabilistic models beat GANs on medical images. arXiv:2212.07501 [eess.IV], 2022.
  7. Pinaya WHL, Tudosiu P, Dafflon J et al.  Brain imaging generation with latent diffusion models. arXiv:2209.07162 [eess.IV], 2022.
  8. Kazerouni A, Aghdam E, Heidari M et al.  Diffusion models for medical image analysis: a comprehensive survey. arXiv:2210.08402 [cs.CV], 2023.
  9. Rombach R, Blattmann A, Lorenz D et al.  High-resolution image synthesis with latent diffusion models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2021:10674–85.
  10. Ruiz N, Li Y, Jampani V et al.  Fine tuning text-to-image diffusion models for subject-driven generation. arXiv:2208.12242 [cs.CV], 2022
  11. Khaled R, Helal M, Alfarghaly O et al.  Categorized contrast enhanced mammography dataset for diagnostic and artificial intelligence research. Sci Data 2022;9:122.
Marketing

 

 

How effective are AI-generated images in marketing?

In the digital age, marketing is shifting from broad dissemination to personalized and precisely targeted strategies (Kim et al. 2022). Personalized marketing aims to recognize and address the individual needs of customers, including tailored emails, customized websites through cookies, recommendation engines, social media engagement, and finely tuned customer services (Dodds, 2024). By delivering personalized messages to their customers, companies can build deeper connections, increase engagement, conversion rates, and loyalty (Hovespian, 2024). At the core of massive personalization in modern marketing are AI systems that go beyond traditional data analysis, processing excessive amounts of user data to adapt strategies in real time to the behavior of each user (Akilkahkov, 2024).

Generative AI models have certainly earned their place in the industry by creating realistic images that appeal to customers, including in the media (Davenport and Mittal, 2022). But how effective are they? A study by TU Munich generated more than ten thousand images with seven modern generative text-to-image models and collected 254,400 human ratings (Hartmann et al., 2024). Respondents rated four of the seven tested models higher for “image quality,” with all seven models performing significantly better than human-created images. Considering the enormous cost difference between human freelancers and generative models, the clear advantage of the models becomes even more apparent: the researchers had to pay $100 for a single image from human freelancers, while with the same budget they could generate 2,500 images using DALL-E 3, the generative model with the highest scores.

Image generation process of the study by Hartmann et al.: A human-made image was converted into text, which was then fed into generative text-to-image models to create synthetic images.

Ethical concerns remain, however, as such generative AI models can be used to circumvent copyrights on training images (Avey, 2023). For example, if we follow the approach described above—converting a copyrighted image into a text representation, then using generative text-to-image models to transform the text into another image and using the synthetic image for marketing purposes—the original creator would practically never be able to detect the misuse of their creative work. Avey warns that most AI models take images from the internet without regard for whether they are copyrighted or not, as stated in the aforementioned report.

 
 
 
 

REFERENCES
  1. Kim, J. (Jay), Kim, T., Wojdynski, B. W., & Jun, H. (2022). Getting a little too personal? positive and negative effects of personalized advertising on online multitaskers. Telematics and Informatics, 71, 101831. https://doi.org/10.1016/j.tele.2022.101831 

  2. Dodds, D. (2024, August 13). Council post: Personalization in marketing: Beyond the buzzword to business impact. Forbes. https://www.forbes.com/councils/forbesagencycouncil/2024/02/27/personalization-in-marketing-beyond-the-buzzword-to-business-impact/ 
  3. Hovsepian, T. (2024, August 13). Council post: The power of personalization: Crafting tailored marketing campaigns for maximum impact. Forbes. https://www.forbes.com/councils/forbesbusinesscouncil/2024/07/22/the-power-of-personalization-crafting-tailored-marketing-campaigns-for-maximum-impact/ 
  4. Akilkhanov, A. (2024, August 13). Council post: AI and personalization in marketing. Forbes. https://www.forbes.com/councils/forbescommunicationscouncil/2024/01/05/ai-and-personalization-in-marketing/ 
  5. Davenport, T. H., & Mittal, N. (2023, August 15). How generative AI is changing creative work. Harvard Business Review. https://hbr.org/2022/11/how-generative-ai-is-changing-creative-work 
  6. Hartmann, J., Exner, Y., & Domdey, S. (2024). The power of Generative Marketing: Can generative AI create Superhuman Visual Marketing Content? International Journal of Research in Marketing. https://doi.org/10.1016/j.ijresmar.2024.09.002 
  7. Avey, C. (2023, December 11). Ethical pros and cons of AI Image Generation. IEEE Computer Society. https://www.computer.org/publications/tech-news/community-voices/ethics-of-ai-image-generation
 
Geology

 

 

How can synthetic images from generative AI models help in modeling heterogeneous rocks?

Pore structure refers to the general characteristics of the size, shape, distribution, and connectivity of the pore space (Jiang et al., 2007). Understanding its features and their impact on the physical properties of rocks, such as permeability, elasticity, and electrical properties, is important in many subfields of geosciences and petroleum engineering (Zhu et al., 2022). However, accurately characterizing the complexity of pore structure remains a challenge (Li et al., 2022), at least partly due to the heterogeneity of rocks, which makes visual examination in laboratory experiments difficult (Sun et al., 2017).

Digital rock modeling is an important method in geology for studying the microstructure and properties of rocks (Fang et al., 2020), which can be subdivided into the following subfields:

  • Digital rock physics (DRP) can be used for the direct quantification of the structural and morphological parameters of rocks and for predicting flow properties at the pore scale (Sadeghnejad et al., 2021). DRP and its non-destructive methods have already become an important complementary method for reservoir characterization three decades ago (Blunt and King, 1991).
  • Digital rock chemistry (DRC) is applied when changes in the pore structure are caused by interactions with dissolved substances (Sadeghnejad et al., 2021).
  • Digital rock biology (DRB) is applied when changes in the pore structure are caused by microbial activities (Sadeghnejad et al., 2021).

Dense sandstone is a heterogeneous rock with diverse mineral compositions and multiscale pore structures, making it difficult to analyze using digital rock modeling (Chi et al., 2024). Chi et al. published in 2024 an attention-guided generative adversarial network that combines X-ray micro-computed tomography (Micro-CT) and scanning electron microscopy (SEM) images to create large-scale, high-precision rock images. Micro-CT images are non-destructive and cheaper but have lower resolution than the destructive and expensive SEM images. The goal was to develop a model capable of generating high-resolution synthetic SEM images from low-resolution Micro-CT images. Although they were not the first to use generative AI models to create high-quality rock images for digital rock physics (Niu et al., 2020; Chen et al., 2020), nor the first to combine Micro-CT and SEM images with generative AI models (Liu and Mukerji, 2022), they also addressed microstructures, including clay morphology.

Schematic representation of the attention-based GAN architecture by Chi et al.

 

The model is based on CycleGAN, where the model aims to improve low-resolution (LR) images to high-resolution (HR) while the adversary tries to reduce the resolution back to LR (Zhu et al., 2017). The authors used two mask generators, a content and an attention mask generator, in which higher-resolution images can be achieved by combining attention masks with their respective content masks and applying them to low-resolution images. The masks have the task of masking all parts of the image except for the small part that we want to draw the model's attention to, allowing the model to learn several characteristic features of different rock components simultaneously.

Comparison between original images and synthetic images. The predicted SEM images in (b) show that the model was able to generate high-resolution SEM images from low-resolution micro-CT images that are very similar to the real SEM images.

 

 
As intended by the authors, the model was able to generate synthetic high-resolution TEM images that preserve the micropores and trace minerals. A model that can create high-resolution synthetic images based on low-resolution real images is likely to provide effective technical support for digital rock modeling, characterization of pore structure, and numerical simulation of rock physics.

 

 
 

REFERENCES
  1. Jiang Z, Wu K, Couples G, Van Dijke MIJ, Sorbie KS, Ma J (2007) Efficient extraction of networks from three-dimensional porous media. Water Resour Res 43(12):W12S03
  2. Zhu LQ, Ma YS, Cai JC, Zhang CM, Wu SG, Zhou XQ (2022) Key factors of marine shale conductivity in southern China-Part II: the influence of pore system and the development direction of shale gas saturation models. J Petrol Sci Eng 209:109516
  3. Li, Xiaobin, Wei, W., Wang, L., Ding, P., Zhu, L., & Cai, J. (2022). A new method for evaluating the pore structure complexity of digital rocks based on the relative value of fractal dimension. Marine and Petroleum Geology, 141, 105694. https://doi.org/10.1016/j.marpetgeo.2022.105694 
  4. Sun, H., Vega, S., & Tao, G. (2017). Analysis of heterogeneity and permeability anisotropy in carbonate rock samples using Digital Rock Physics. Journal of Petroleum Science and Engineering, 156, 419–429. https://doi.org/10.1016/j.petrol.2017.06.002 
  5. Fang, H.-H., Sang, S.-X., & Liu, S.-Q. (2020). Three-dimensional spatial structure of the macro-pores and flow simulation in anthracite coal based on X-ray μ-CT scanning data. Petroleum Science, 17(5), 1221–1236. https://doi.org/10.1007/s12182-020-00485-3
  6. Sadeghnejad, S., Enzmann, F., & Kersten, M. (2021). Digital Rock Physics, Chemistry, and biology: Challenges and prospects of pore-scale modelling approach. Applied Geochemistry, 131, 105028. https://doi.org/10.1016/j.apgeochem.2021.105028 
  7. Blunt, M., and King, P. (1991). Relative Permeabilities from Two- and Three-Dimensional Pore-Scale Network Modelling. Transport Porous Med. 6, 407–433. doi:10.1007/bf00136349
  8. Chi, P., Sun, J., Yan, W., & Luo, X. (2024). Multiscale fusion of tight sandstone digital rocks using attention-guided generative Adversarial Network. Marine and Petroleum Geology, 160, 106647. https://doi.org/10.1016/j.marpetgeo.2023.106647 
  9. Niu, Y., Wang, Y. D., Mostaghimi, P., Swietojanski, P., & Armstrong, R. T. (2020). An innovative application of generative adversarial networks for physically accurate rock images with an unprecedented field of view. Geophysical Research Letters, 47(23). https://doi.org/10.1029/2020gl089029 
  10. Chen, H., He, X., Teng, Q., Sheriff, R. E., Feng, J., & Xiong, S. (2020). Super-resolution of real-world rock microcomputed tomography images using cycle-consistent generative adversarial networks. Physical Review E, 101(2). https://doi.org/10.1103/physreve.101.023305 
  11. Liu, M., & Mukerji, T. (2022). Multiscale fusion of digital rock images based on deep generative adversarial networks. Geophysical Research Letters, 49(9). https://doi.org/10.1029/2022gl098342 
  12. Zhu, J.-Y., Park, T., Isola, P., & Efros, A. A. (2017). Unpaired image-to-image translation using cycle-consistent adversarial networks. 2017 IEEE International Conference on Computer Vision (ICCV), 2242–2251. https://doi.org/10.1109/iccv.2017.244 
 
 

!     Consider what other tasks in your field could be supported by generative AI models.