Across Domains

Supervised learning "Across Domains"

 

Supervised learning models are used to support professionals from different fields in performing various tasks.

Below, select one of the areas closest to your area of expertise and explore different tasks that can be supported by supervised learning models:

 

Medicine
 

The supervised learning approach is being used more and more frequently in the medical field. Below are some examples of how the supervised learning approach is used in a medical context.

How can I predict the most suitable dose of medication for my patient?

It can be difficult for healthcare professionals to predict the correct dose of medication for a particular patient. In some cases, incorrect dosing can result in devastating side effects.

Some studies have investigated the use of supervised learning to create predictive dosing models that can assist healthcare professionals in making decisions. For example, Ahn et al. developed warfarin dose prediction models using supervised learning techniques to assist healthcare professionals in determining warfarin dosing.

How can I predict illnesses in patients in good time?

This is another example. There are diseases that are very dangerous and require very early action. Therefore, timely identification and detection of a certain type of disease and treatment is crucial to avoid further deterioration of the patient's health.

This is a challenge where supervised learning can also be used to help healthcare professionals recognize diseases earlier (i.e. provide clinical decision support). Information about previous patients can be made available to supervised learning algorithms to develop predictive models to help healthcare professionals determine the likelihood of a patient suffering from a disease or serious condition. 

For example, Menni et al. have developed a model that can predict whether a user has COVID-19 based on their app data, specifically the symptoms they report. Kim, Cho and Oh have developed a model to help healthcare professionals diagnose glaucoma based on patient information such as age, eye pressure, corneal thickness and others. 

Final thoughts

The use of AI systems in routine clinical care today represents an important but still largely untapped opportunity, as the medical AI community must overcome the complex ethical, technical and human-centered challenges required for safe and effective use (Rajpurkar et al., 2022). The examples shown here should NOT be used by physicians and patients to make treatment decisions, as they are purely educational and research resources.


REFERENCES
  • Rajpurkar, P., E. Chen, O. Banerjee, and E.J. Topol. “AI in Health and Medicine.” Nature Medicine 28, no. 1 (2022): 31–38. https://doi.org/10.1038/s41591-021-01614-0.
  • Roy, S., T. Meena, and S.-J. Lim. “Demystifying Supervised Learning in Healthcare 4.0: A New Reality of Transforming Diagnostic Medicine.” Diagnostics 12, no. 10 (2022). https://doi.org/10.3390/diagnostics12102549.
  • Menni, Cristina, Ana M. Valdes, Maxim B. Freidin, Carole H. Sudre, Long H. Nguyen, David A. Drew, Sajaysurya Ganesh, et al. “Real-Time Tracking of Self-Reported Symptoms to Predict Potential COVID-19.” Nature Medicine 26, no. 7 (July 2020): 1037–40. https://doi.org/10.1038/s41591-020-0916-2.
  • Kim, Seong Jae, Kyong Jin Cho, and Sejong Oh. “Development of Machine Learning Models for Diagnosis of Glaucoma.” PLOS ONE 12, no. 5 (May 23, 2017): e0177726. https://doi.org/10.1371/journal.pone.0177726.
  • Ahn, S. “Building and Analyzing Machine Learning-Based Warfarin Dose Prediction Models Using Scikit-Learn.” Translational and Clinical Pharmacology 30, no. 4 (2022): 172–81. https://doi.org/10.12793/tcp.2022.30.e22.
 
Biology

 

Supervised learning is increasingly being used in the biological field. Below we see an example of how this approach is applied in a biological context.

 

How can we predict whether a novel drug will bind to a protein?

Many research projects in the field of biology aim to develop drugs for diseases that were previously difficult to cure. Since proteins are at the center of biological processes, it is undisputed that new drug discoveries are made by targeting proteins, especially in rare diseases where a specific protein is responsible for the underlying biological process. However, isolating proteins to measure their affinity for a drug is a lengthy process, as there is not just one, but many possible candidates for the drug.

From the many drugs already known to interact with specific proteins, we can train a machine learning model to predict the affinity for a new drug-protein pair - a problem known as DTI (drug-target interaction). Sound familiar? Supervised learning can leverage the vast protein and drug databases to facilitate the drug discovery process. The difficult question here is how to select the features for prediction. Guvenilir et al. built prediction models based on different features of a protein and compared their performance. One such feature was the amino acid composition, which was calculated separately for each amino acid, while more complex features were also used. The design and selection of features for an AI model is a very important step in AI development and a phase where the expertise of experts in the field - people like you - is most needed.

 

How can we predict the structure of a protein from its sequences?

The 3D conformation of the amino acids and their functional groups plays a central role in determining the activity of the protein. Clearly, knowledge of protein structure is the key to the most accurate predictions of protein-drug activity.

As there are far more known protein sequences than their structures, international competitions have been launched to test computational prediction methods, called CASP (Critical Assessment of Structure Prediction), with the aim of bridging this information gap between sequence and structure.

While it is always interesting to explore new approaches presented at CASP, here we would like to introduce AlphaFold, which suddenly emerged as the winner in 2020 and nearly perfected its prediction capability by 2022. Like many other competitors, it used supervised learning, but not directly on sequence-to-structure relationships, but first on sequence-to-sequence relationships to identify sequences with known structures that are evolutionarily homologous to the target. This was another example of how beneficial it can be to integrate different data for predictions. AlphaFold was awarded the 2024 Nobel Prize in Chemistry.

 
 
 
Final thoughts

Life is a complex phenomenon and cannot be described with simple numbers. Therefore, we have to choose very carefully what we want to predict and what information to base the prediction on. This means that more biological knowledge and intuition will be required to develop increasingly advanced AI models. If we then use these models to gain more knowledge and insight, we will be one step closer to fully understanding who we are and how we live.
 
 
REFERENCES
  1. Atas Guvenilir, H., Doğan, T. How to approach machine learning-based prediction of drug/compound–target interactions. J Cheminform 15, 16 (2023). https://doi.org/10.1186/s13321-023-00689-w
  2. Thomas Luechtefeld, Dan Marsh, Craig Rowlands, Thomas Hartung, Machine Learning of Toxicological Big Data Enables Read-Across Structure Activity Relationships (RASAR) Outperforming Animal Test Reproducibility, Toxicological Sciences, Volume 165, Issue 1, September 2018, Pages 198–212, https://doi.org/10.1093/toxsci/kfy152
  3. Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). https://doi.org/10.1038/s41586-021-03819-2
Chemistry

 

Supervised learning is increasingly being used in chemistry. Below are some examples of how these techniques are used in chemical research.

 

How can we predict the chemical properties of a substance from other chemical properties?

Although we can measure properties experimentally, this sometimes requires more budget and time than we have available. Therefore, chemists are trying to answer a question known as Quantitative Property-Property Relation (QPPR).

QPPR involves identifying numerical relationships between chemical properties that allow one unknown property to be predicted from others. This not only facilitates the experimental measurement process, but also helps us to better understand the “chemical” interplay behind these properties.

 

How can we predict the chemical properties of a substance from its structure?

For QPPR, many other properties must be measured to predict a property. An experienced chemist can predict a property qualitatively just by looking at the structure. Quantitative Structure-Activity Relation (QSAR) models aim to do this quantitatively.

If we could accurately predict all chemical properties from structure alone, this would greatly accelerate chemical research. To date, however, there is still no “perfect” QSAR model, which is why QSAR models remain an active area of research. Due to their versatility in chemical research, many start-ups and institutions are trying to develop better solutions to this problem.

Here we present a rare example that is not hidden behind a paywall: OPERA, an open-source QSAR model developed by K. Mansouri et al. It can predict a variety of chemical properties, including pH. Surprisingly, this model is mainly based on a simple supervised learning algorithm called k-nearest neighbors (KNN). OPERA has a graphical user interface (GUI), so it can be used without programming knowledge - proving that even people who don't want to learn programming should familiarize themselves with AI techniques.

 
 

Final thoughts 

Observations and the attempt to explain them have always been at the center of science. Machine learning algorithms, including the supervised learning approach, are undoubtedly capable of recognizing patterns and deriving precise predictions from them. It is no wonder that they are increasingly being used in various fields of science. If we use them not only to predict but also to gain knowledge, we can gain valuable insights that were previously hidden from our naked eye.



REFERENCES
  1. L. Deborah et al. “A Gentle Introduction to Machine Learning for Chemists: An Undergraduate Workshop Using Python Notebooks for Visualization, Data Processing, Analysis, and Modeling.” Journal of Chemical Education 98, no. 9 (September 14, 2021): 2892–98. https://doi.org/10.1021/acs.jchemed.1c00142.
  2. DeJongh, J., Verhaar, H. & Hermens, J. A quantitative property-property relationship (QPPR) approach to estimate in vitro tissue-blood partition coefficients of organic chemicals in rats and humans. Arch Toxicol 72, 17–25 (1997). https://doi.org/10.1007/s002040050463
  3. Mansouri, K., Grulke, C.M., Judson, R.S. et al. OPERA models for predicting physicochemical properties and environmental fate endpoints. J Cheminform 10, 10 (2018). https://doi.org/10.1186/s13321-018-0263-1
 
Economics
 

 

The supervised learning approach is increasingly being adopted in the field of economics. In the following section, we will examine some examples of how the supervised learning approach has been applied in the context of economics.

How can supervised machine learning impact the buying and selling of securities such as stocks and options?

Stock and options traders rely on their experience to analyze the prices of such securities. It takes a lot of time and effort to learn the skills to deduce patterns in the financial market. Their conclusions are often limited by the amount of data humans can process. This can be overcome by supervised learning.

Just like highly skilled traders, supervised machine learning learns from past trends, but unlike traders, computers can gather much more data to create models. And so they can generate a more accurate model. This can significantly improve the success rate in securities trading.

For example, in Bazrkar and Hosseini, 2022, the authors explore the possibility of using various machine learning techniques to predict stock prices. The study concludes that a support vector machine (SVM) can be used to predict stock prices with an accuracy of over 90%.

 
 

How can supervised machine learning be used to detect fraud and money laundering?

Fraudulent insurance claims are a major problem for insurance companies. Detecting fraudulent claims can be time-consuming and labor-intensive, and reviewing each individual case is cumbersome. In such cases, supervised learning can be used as a complementary method to speed up the process. In Debener, Heinke and Kriebel, 2023, the authors discuss the application of supervised learning for fraud detection.

Similarly, detecting cases of money laundering can be tedious. Given the volume of financial transactions, it can be incredibly difficult to detect a case of money laundering. Even in such cases, supervised machine learning can help. Authors Alsuwailem, Salem and Saudagar, 2023, discuss one such application. The article describes the use case of supervised learning to detect cases of money laundering in Saudi Arabia. The Financial Intelligence Unit of Saudi Arabia has successfully used SML to detect and reduce financial crimes with an accuracy of up to 93%.

 
 

How can a recession be predicted with supervised machine learning?

A recession can have devastating consequences for any economy. It often leads to unemployment and inflation, both of which worsen the normal standard of living of citizens. However, a recession is difficult to predict.

In such a scenario, supervised learning can be used. Supervised learning algorithms can use data from previous recessions to create models to predict a recession. For example, researchers at Malladi used supervised machine learning in 2022 to predict a post-COVID-19 recession.

 
 
 
REFERENCES
  1. Mohammad Javad Bazrkar and Soodeh Hosseini. 2022. Predict Stock Prices Using Supervised Learning Algorithms and Particle Swarm Optimization Algorithm. Comput. Econ. 62, 1 (Jun 2023), 165–186. https://doi.org/10.1007/s10614-022-10273-3
  2. Wu, Mu-En, Jia-Hao Syu, and Chien-Ming Chen. “Kelly-Based Options Trading Strategies on Settlement Date via Supervised Learning Algorithms.” Computational Economics 59, no. 4 (April 1, 2022): 1627–44. https://doi.org/10.1007/s10614-021-10226-2.
  3. Debener, Jörn, Volker Heinke, and Johannes Kriebel. “Detecting Insurance Fraud Using Supervised and Unsupervised Machine Learning.” Journal of Risk and Insurance 90, no. 3 (2023): 743–68. https://doi.org/10.1111/jori.12427.
  4. Alsuwailem, Alhanouf Abdulrahman Saleh, Emad Salem, and Abdul Khader Jilani Saudagar. “Performance of Different Machine Learning Algorithms in Detecting Financial Fraud.” Computational Economics 62, no. 4 (December 1, 2023): 1631–67. https://doi.org/10.1007/s10614-022-10314-x.
  5. Malladi, Rama K. “Application of Supervised Machine Learning Techniques to Forecast the COVID-19 U.S. Recession and Stock Market Crash.” Computational Economics, October 26, 2022. https://doi.org/10.1007/s10614-022-10333-8.
  6. Clithero, John A., Jae Joon Lee, and Joshua Tasoff. “Supervised Machine Learning for Eliciting Individual Demand.” American Economic Journal: Microeconomics 15, no. 4 (November 2023): 146–82. https://doi.org/10.1257/mic.20210069.
 
Geosciences
 

Supervised learning is increasingly being used in the geosciences. Below are some examples of how supervised learning is used in geoscience research.

How can I predict a meteorological property based on other properties?

Although meteorology has a direct influence on many aspects of our lives, it is often difficult to collect meteorological data directly from our own neighborhood. Until now, we have had to deduce the conditions in our surroundings from a few nearby measuring stations. For example, if you live in Treptow-Köpenick in Berlin, the weather forecasts on TV will tell you about the weather in Berlin, but not directly about Treptow-Köpenick. So we had to assume that the weather in Treptow is almost the same as in Berlin, even though the measurements were actually taken in Tempelhof or Schönefeld. Even online weather platforms that support regional forecasts have to derive the data from nearby stations.

 

Geology-Fig. 1. Screenshot of “Windfinder”, an online platform for current wind data and forecasts. There is no weather station in Köpenick that is connected to “Windfinder”, so the information has to be derived from nearby stations.

There is no doubt that air quality is of great importance to living standards, but air quality varies greatly from street to street and block to block, which can make such inferences from nearby stations inaccurate. Zheng et al. from Microsoft Research in Beijing used a semi-supervised learning approach to derive air quality information with an accuracy of over 80% by incorporating various factors such as street maps, traffic and weather information. This demonstrates the potential of AI systems, including supervised learning algorithms, to draw accurate conclusions from ‘big data’ - something that would be too time-consuming for a human to process in real time.



How can I predict a meteorological property for the future based on past records?

The next natural question is how to predict the weather for tomorrow. Weather forecasts are always one of the most exciting parts of the news, aren't they? Machine learning algorithms, including supervised learning algorithms, are continuously improving our weather forecasts. There are even weather forecasts based solely on machine learning that are available online. The Karlsruhe Institute of Technology (KIT) runs its own website where weather forecasts developed by machine learning algorithms are published: KIT-Wetter. It might be interesting to take a look at the weather forecast for tomorrow and see for yourself how well machine learning can handle this task.

 

Final thoughts
 
We all know the situation: the news says it won't rain tomorrow, but then it does - or vice versa. Sometimes they say it's raining in my city, but when I look out of the window, it's not. With the new machine learning algorithms, these false predictions could soon be a thing of the past. Nevertheless, as inaccurate as the forecasts were in the past, the weather data collected over decades from all over the world offers supervised learning algorithms an enormous learning basis. The more mistakes we make and the more experience we gain as a human race, the fewer mistakes we could make in future thanks to our AI-supported learning partners.



REFERENCES
  1. Deutscher Wetterdienst. (n.d.). Metanavigation. Wetter und Klima - Deutscher Wetterdienst - Our services - Climate at selected weather stations in Berlin and Brandenburg. https://www.dwd.de/EN/ourservices/cos/berlin_brandenburg.html
  2. Hartmann, K., Krois, J., Rudolph, A. (2023): Statistics and Geodata Analysis using R (SOGA-R). Department of Earth Sciences, Freie Universitaet Berlin. https://www.geo.fu-berlin.de/en/v/soga-r/index.html
  3. Yu Zheng, Furui Liu, and Hsun-Ping Hsieh. 2013. U-Air: when urban air quality inference meets big data. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining (KDD '13). Association for Computing Machinery, New York, NY, USA, 1436–1444. https://doi.org/10.1145/2487575.2488188
 
 

!

Think about what other tasks in your field of study could use supervised learning models.