2. Transparency of AI-systems

Blackbox-Models "Across Domains"

Now that we have a better idea of how problematic black-box models can be, it's a good time to present some examples of problems and challenges associated with black-box models in various fields.

Select a field below that is closest to your area of expertise and explore various problems in different domains where black-box models may raise concerns.

 
 
 
Medicine

Despite the promising possibilities of AI in healthcare, its potential remains untapped (Kelly et al., 2019). In one study, images of benign moles were diagnosed as malignant when the images were simply rotated or a small amount of noise was added (Finlayson et al., 2019). Although these errors may seem minor at first glance, they are difficult to fix because developers have little or no control over the internal parameters (weights) of the AI model due to its black-box nature (Zhai et al., 2023). Although such manipulations (e.g., rotations and noise) were artificial in the research context, it would be difficult to use such a sensitive model in the real world.

Since black-box models can make mistakes that are impossible to interpret, and human lives may depend on them, special care should be taken to make AI models in healthcare as transparent and accountable as possible. Apart from this issue, a review by Ciobanu-Caraus et al. (2024) found that half of the published machine learning studies in healthcare were not based on publicly available data, and only one-fifth of them shared their code. More than three-quarters of the studies used data from only one institution, which ignores the risk of bias.

Literature

  1. Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G., & King, D. (2019). Key challenges for delivering clinical impact with artificial intelligence. BMC medicine, 17(1), 195. https://doi.org/10.1186/s12916-019-1426-2
  2. Finlayson, S. G., Bowers, J. D., Ito, J., Zittrain, J. L., Beam, A. L., & Kohane, I. S. (2019). Adversarial attacks on medical machine learning. Science (New York, N.Y.), 363(6433), 1287–1289. https://doi.org/10.1126/science.aaw4399
  3. Zhai, Z., Li, P. & Feng, S. State of the art on adversarial attacks and defenses in graphs. Neural Comput & Applic 35, 18851–18872 (2023). https://doi.org/10.1007/s00521-023-08839-9
  4. Ciobanu-Caraus, O., Aicher, A., Kernbach, J. M., Regli, L., Serra, C., & Staartjes, V. E. (2024). A critical moment in machine learning in medicine: on reproducible and interpretable learning. Acta neurochirurgica, 166(1), 14. https://doi.org/10.1007/s00701-024-05892-8

 
Economy
 

Apple Card, a credit service operated by Goldman Sachs, reportedly offers different credit limits based on gender, even though the financial profile is similar (Vigdor 2019). Bloomberg reported a striking example (Natarajan and Nasiripour, 2019) in which a husband was given a credit limit 20 times higher than his wife, even though his wife actually had a higher credit rating. The BBC suspected that the algorithm behind Apple Card was trained with biased data suggesting that women pose a greater financial risk than men, and accused it of being a “black box” that does not reveal how credit scores are determined (BBC, 2019).

A similar case was reported with Amazon's AI model, which was used to automatically hire employees and also discriminated against women compared to men (Dastin 2018). After a long search, the engineers discovered that the cause of the problem lay in the training data, which came from the company's survey conducted over a period of 10 years.

The black-box nature of some AI models, such as DNNs, makes it difficult to solve such a problem, as it is difficult to know which aspect or feature of the input led the model to certain results, and it is even more difficult to adjust the internal parameters of the model to correct such behavior.

 
Literature
  1. Vigdor, N. (2019, November 10). Apple card investigated after gender discrimination complaints. The New York Times. https://www.nytimes.com/2019/11/10/business/Apple-credit-card-investigation.html 
  2. Natarajan, S., & Nasiripour, S. (2019, November 9). Goldman Sachs probed after viral tweet about Apple Card. Bloomberg.com. https://www.bloomberg.com/news/articles/2019-11-09/viral-tweet-about-apple-card-leads-to-probe-into-goldman-sachs 
  3. BBC. (2019, November 11). Apple’s “sexist” credit card investigated by US Regulator. BBC News. https://www.bbc.com/news/business-50365609 
  4. Dastin, J. (2018, October 11). Insight - Amazon scraps secret AI recruiting tool that showed bias against women | reuters. Reuters. https://www.reuters.com/article/world/insight-amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women-idUSKCN1MK0AG/ 
  5. Ribeiro, M., Singh, S., & Guestrin, C. (2016). “Why should i trust you?”: Explaining the predictions of any classifier. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. https://doi.org/10.18653/v1/n16-3020 
Chemistry

The 2018 NOMAD Kaggle competition focused on models for predicting the formation energies of transparent conductive oxides (TCOs) (Sutton et al., 2019). Three participants achieved leading results with very small margins: the n-gram method, which was adopted from the field of natural language processing (NLP) (Sutton et al., 2019), Smooth Overlap of Atomic Positions (SOAP) (Bartók et al., 2010, Bartók et al., 2013), and the Many-Body Tensor Representation (MBTR) (Huo and Rupp, 2022). Although all three methods proved unsuitable for screening purposes in practice, the MBTR method was found to be applicable to a subset of materials identified by a subgroup discovery (SGD) study by Sutton et al. (2020). In other words, n-gram and SOAP were equally unsuitable for different materials, while MBTR performed acceptably well in a specific material range but performed worse in others.

Unless specifically designed otherwise, deep neural networks (DNNs) treat all inputs equally—that is, the loss of one input is not favored over the loss of another (Fernando and Tsokos, 2022). This can lead to models being trained and tested on a large amount of data but evaluated on a single value. This can disadvantage models that work very well for a subset of data (Boley et al., 2020). As an analogy, it would make more sense to give a student separate grades for different courses rather than assigning them a single overall grade for the entire semester. With a single grade for the entire semester, it becomes more difficult for students to recognize what they are good at and what they still need to improve. Ultimately, this is why people choose a field of study and a profession. Just like people, models are not equally good at all tasks, but have strengths and weaknesses (Jha et al., 2018).

The big difference, however, is that we know how to divide the “semester grade” into different courses by simply assigning different grades for different courses. However, this is not always possible for AI models because, in technical terms, the features used by AI models are not necessarily compatible with the features used by humans (Zhang et al., 2024). Therefore, it is difficult to select or design the best model for analyzing a specific material.

This illustrates the difficulty of applying black-box models in general to a large group of inputs. This problem can be addressed by identifying the strengths and weaknesses of the models. Nevertheless, problems remain unsolved because we do not know why the model is better in some areas and worse in others. We conclude this discussion with a related quote from a famous psychiatrist:

“The shoe that fits one person may pinch another; there is no recipe for life that is suitable for all cases.” - Karl Jung

Literature
  1. Sutton, C., Ghiringhelli, L.M., Yamamoto, T. et al. Crowd-sourcing materials-science challenges with the NOMAD 2018 Kaggle competition. npj Comput Mater 5, 111 (2019). https://doi.org/10.1038/s41524-019-0239-3
  2. Bartók, A. P., Payne, M. C., Kondor, R., & Csányi, G. (2010). Gaussian approximation potentials: The accuracy of quantum mechanics, without the electrons. Physical Review Letters, 104(13). https://doi.org/10.1103/physrevlett.104.136403 
  3. Bartók, A. P., Kondor, R., & Csányi, G. (2013). On representing Chemical Environments. Physical Review B, 87(18). https://doi.org/10.1103/physrevb.87.184115 
  4. Huo, H., & Rupp, M. (2022). Unified representation of molecules and crystals for machine learning. Machine Learning: Science and Technology, 3(4), 045017. https://doi.org/10.1088/2632-2153/aca005 
  5. K. R. M. Fernando and C. P. Tsokos, "Dynamically Weighted Balanced Loss: Class Imbalanced Learning and Confidence Calibration of Deep Neural Networks," in IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 7, pp. 2940-2951, July 2022, doi: 10.1109/TNNLS.2020.3047335.
  6. Sutton, C., Boley, M., Ghiringhelli, L.M. et al. Identifying domains of applicability of machine learning models for materials science. Nat Commun 11, 4428 (2020). https://doi.org/10.1038/s41467-020-17112-9
  7. Jha, D., Ward, L., Paul, A. et al. ElemNet: Deep Learning the Chemistry of Materials From Only Elemental Composition. Sci Rep 8, 17593 (2018). https://doi.org/10.1038/s41598-018-35934-y
  8. Zhang, S., Xia, B., Zhang, Z., Wu, Q., Sun, F., Hu, Z., & Sun, Y. (2024). Automated Molecular Concept Generation and Labeling with Large Language Models. ArXiv, abs/2406.09612.
 


!
What sensitive issues in your field of study could raise concerns when black box models are used? Think about it or search the internet for an example.