CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
Marko ŘEHÁČEK, Vítězslav DUŠEK, Martin RUSINKO, Vít NOVÁČEK
CIKM 2026 · DOI ↗
Names in bold are current members and collaborators of the group.
Marko ŘEHÁČEK, Vítězslav DUŠEK, Martin RUSINKO, Vít NOVÁČEK
CIKM 2026 · DOI ↗
Radovan TOMÁŠIK, Tobias KUSSEL, Zdenka DUDOVÁ, Radoslava KACOVÁ, Roman HRSTKA, Martin LABLANS, Petr HOLUB
BMC Medical Informatics and Decision Making · DOI ↗
Assessing data quality in federated health data systems presents unique challenges, particularly when data custodians cannot expose raw data due to privacy regulations. Traditional quality assessment approaches often require centralised access, which conflicts with the principles of data sovereignty and confidentiality. In this study, we evaluate the utility of federated data quality assessment with differential privacy techniques to safeguard sensitive health data. The aim is to develop tooling and demonstrate a proof-of-concept implementation over a synthetic dataset of observational medical data. We present a privacy-preserving framework for evaluating data quality in federated environments using differential privacy. Our approach enables individual data providers to compute local quality metrics and share only aggregated, privacy-protected results. We implement a proof-of-concept that supports predefined quality checks across different data models and demonstrate how meaningful insights into data quality can be obtained without compromising sensitive information. This work demonstrates that differential privacy can be effectively applied to enable federated quality assessment in health data networks without compromising individual privacy. By implementing a proof-of-concept system over synthetic health data, we show that it is possible to obtain meaningful quality metrics in a decentralised setting.
Petr ZELINA, Marko ŘEHÁČEK, Jana HALÁMKOVÁ, Lucia BOHOVICOVÁ, Martin RUSINKO, Vít NOVÁČEK
TSD 2025 (Springer LNCS) · DOI ↗ · PDF ↗
Clinical notes hold rich yet unstructured details about diagnoses, treatments, and outcomes that are vital to precision medicine but hard to exploit at scale. We introduce a method that represents each patient as a matrix built from aggregated embeddings of all their notes, enabling robust patient similarity computation based on their latent low-rank representations. Using clinical notes of 4,267 Czech breast-cancer patients and expert similarity labels from Masaryk Memorial Cancer Institute, we evaluate several matrix-based similarity measures and analyze their strengths and limitations across different similarity facets, such as clinical history, treatment, and adverse events. The results demonstrate the usefulness of the presented method for downstream tasks, such as personalized therapy recommendations or toxicity warnings.
Radoslava KACOVÁ, Tomáš HOUFEK, Ondřej HORKÝ, Radovan TOMÁŠIK, Jan KURÁŇ, Michal RŮŽIČKA, Roman HRSTKA, Vít NOVÁČEK, Zdenka DUDOVÁ
Data Intelligence · DOI ↗
In the dynamic environment of hospitals, valuable real-world data often remain underutilised despite their potential to revolutionize cancer research and personalised medicine. This study explores the challenges and opportunities in managing hospital-generated data, particularly within the Masaryk Memorial Cancer Institute (MMCI) in Brno, Czech Republic. Utilizing Next-Generation Sequencing (NGS) technology, MMCI generates substantial volumes of genomic data. Due to inadequate curation, these data remain difficult to integrate with clinical records for secondary use (such as personalised treatment outcome prediction and patient stratification based on their genomic profiles). This paper proposes solutions based on the FAIR principles (Findability, Accessibility, Interoperability, and Reusability) to enhance data sharing and reuse. The primary output of our work is the development of an automated pipeline that continuously processes and integrates NGS data with clinical and biobank information upon their creation. It stores the data in a special secured repository for sensitive data in a structured form to ensure smooth retrieval.
Mohan TIMILSINA, Samuele BUOSI, Maria TORRENTE, Mariano PROVENCIO, Manuel COBO, Delvys Rodriguez ABREU, Rafael Lopez CASTRO, Enric CARCERENY, Edward CURRY, Vít NOVÁČEK
Expert Systems with Applications · DOI ↗
In this study, we evaluate the effectiveness of foundational artificial intelligence (AI) models, particularly large language models (LLMs), in comparison to traditional machine learning methods for predicting tumor relapse in patients with non-small-cell lung cancer (NSCLC). With a high recurrence risk in NSCLC, early and accurate prediction is essential for improving patient outcomes and guiding treatment decisions. Our analysis utilizes a dataset of 1,348 patients, examining the performance of traditional machine learning models such as Random Forest, alongside cutting-edge LLMs like Mistral-7B, LLaMA-7B, Falcon-7B, and GPT-based models. While the Random Forest model slightly outperforms Mistral-7B in precision–recall for relapse prediction, the comparable results suggest that both approaches offer valuable insights for early relapse detection. This study underscores the potential of integrating classical machine learning with foundational AI models to enhance predictive accuracy in cancer prognosis, providing pathways for more personalized medical interventions.
Radovan TOMÁŠIK, Šimon KOŇÁR, Niina EKLUND, Cacilia ENGELS, Zdenka DUDOVÁ, Radoslava KACOVÁ, Roman HRSTKA, Petr HOLUB
Journal of Biomedical Informatics · DOI ↗
Biobanks and biomolecular resources are increasingly central to data-driven biomedical research, encompassing not only metadata but also granular, sample-related data from diverse sources such as healthcare systems, national registries, and research outputs. However, the lack of a standardised, machine-readable format for representing such data limits interoperability, data reuse and integration into clinical and research environments. While MIABIS provides a conceptual model for biobank data, its abstract nature and reliance on heterogeneous implementations create barriers to practical, scalable adoption. This study presents a pragmatic, operational implementation of MIABIS focused on enabling real-world exchange and integration of sample-level data. We systematically evaluated established data exchange standards, comparing HL7 FHIR and OMOP CDM with respect to their suitability for structuring sample-related data in a semantically robust and machine-readable form. Based on this analysis, we developed a FHIR-based representation of MIABIS that supports complex biobank structures and enables integration with federated data infrastructures. Supporting tools, including a Python library and an implementation guide, were created to ensure usability across diverse research and clinical contexts. We created nine interoperable FHIR profiles covering core MIABIS entities, ensuring consistency with FHIR standards. To support adoption, we developed an open-source Python library that abstracts FHIR interactions and provides schema validation for MIABIS-compliant data. The library was integrated into an ETL tool in operation at Czech Node of BBMRI-ERIC, European Biobanking and Biomolecular Resources Research Infrastructure, to demonstrate usability with real-world sample-related data. Separately, we validated the representation of MIABIS entities at the organisational level by converting the data structures of BBMRI-ERIC Directory into FHIR, demonstrating compatibility with federated data infrastructures. This work delivers a machine-readable, interoperable implementation of MIABIS, enabling the exchange of both organisational and sample-level data across biobanks and health information systems. By integrating MIABIS with HL7 FHIR, we provide a host of reusable tools and mechanisms for further evolution of the data model. Combined, these benefits can help with the integration into clinical and research workflows, supporting data discoverability, reuse, and cross-institutional collaboration in biomedical research.
Petr ZELINA
RASLAN 2024 · PDF ↗
Entity extraction in clinical texts is essential for converting unstructured data in clinical notes into structured formats, facilitating large-scale analysis and clinical decision support. Traditional methods often rely on handcrafted regular expressions (regexes), which, while effective, demand significant time and specialized knowledge to create – resources that healthcare professionals may lack. We introduce a novel approach leveraging large language models (LLMs) to automate regex generation for clinical entity extraction. Our method involves prompting LLMs to generate regex patterns from examples, followed by iterative refinement using a feedback loop. Despite regex limitations, this approach is practical for extracting frequently patterned information common in clinical texts, such as dates, specific data about medical procedures or event detection. Our experiments on Czech clinical notes show this method outperforms current SOTA genetic-programming-based methods for generating regular expression patterns from examples, especially when there are few of them.
Samuele BUOSI, Mohan TIMILSINA, Maria TORRENTE, Mariano PROVENCIO, Dirk FEY, Vít NOVÁČEK
Computers in Biology and Medicine · DOI ↗
The recurrence of low-stage lung cancer poses a challenge due to its unpredictable nature and diverse patient responses to treatments. Personalized care and patient outcomes heavily rely on early relapse identification, yet current predictive models, despite their potential, lack comprehensive genetic data. This inadequacy fuels our research focus-integrating specific genetic information, such as pathway scores, into clinical data. Our aim is to refine machine learning models for more precise relapse prediction in early-stage non-small cell lung cancer. To address the scarcity of genetic data, we employ imputation techniques, leveraging publicly available datasets such as The Cancer Genome Atlas (TCGA), integrating pathway scores into our patient cohort from the Cancer Long Survivor Artificial Intelligence Follow-up (CLARIFY) project. Through the integration of imputed pathway scores from the TCGA dataset with clinical data, our approach achieves notable strides in predicting relapse among a held-out test set of 200 patients. By training machine learning models on enriched knowledge graph data, inclusive of triples derived from pathway score imputation, we achieve a promising precision of 82% and specificity of 91%. These outcomes highlight the potential of our models as supplementary tools within tumour, node, and metastasis (TNM) classification systems, offering improved prognostic capabilities for lung cancer patients. In summary, our research underscores the significance of refining machine learning models for relapse prediction in early-stage non-small cell lung cancer. Our approach, centered on imputing pathway scores and integrating them with clinical data, not only enhances predictive performance but also demonstrates the promising role of machine learning in anticipating relapse and ultimately elevating patient outcomes.
Samuele BUOSI, Mohan TIMILSINA, Adriann JANIK, Luca COSTABELLO, Maria TORRENTE, Mariano PROVENCIO, Dirk FEY, Vít NOVÁČEK
Expert Systems with Applications · DOI ↗
Motivation: Low-stage lung cancer is known to recur unpredictably, and patients receiving various treatment methods like radiation, chemotherapy, and immunotherapies have been seen to respond very differently. Identifying a priori if a patient is going to relapse or not could make a difference in terms of saving lives and personalized care offered. In this work, we provide an answer to the following research question: Is it possible to enhance the machine learning (ML) of the estimated probability of relapse in early-stage non-small-cell lung cancer (NSCLC) patients with aneuploidy imputation scores? Results: To predict recurrence in 1,348 early-stage (I–II) NSCLC patients, we train graph ML models utilizing the Spanish pulmonary cancer group knowledge graph enriched with triples from pathway imputation. ML models trained on Knowledge graph data enriched with triples from pathway score imputation present an 82% Precision and 91% Specificity in predicting relapse over 200 patients from a held-out test set. ML models trained using graphs data could prove useful supplemental tool in the TNM classification systems and improve a lung cancer patient's prognosis.

Petr ZELINA, Jana HALÁMKOVÁ, Vít NOVÁČEK
IEEE Transactions on Nanobioscience · DOI ↗
TLDR: We extend our work on finding structure in unstructured clinical notes with better segment grouping, analysis and ways to map segment types to existing dictionaries.
This work is motivated by the scarcity of tools for accurate, unsupervised information extraction from unstructured clinical notes in computationally underrepresented languages, such as Czech. We introduce a stepping stone to a broad array of downstream tasks such as summarisation or integration of individual patient records, extraction of structured information for national cancer registry reporting or building of semi-structured semantic patient representations that can be used for computing patient embeddings. More specifically, we present a method for unsupervised extraction of semantically-labeled textual segments from clinical notes and test it out on a dataset of Czech breast cancer patients, provided by Masaryk Memorial Cancer Institute (the largest Czech hospital specialising exclusively in oncology). Our goal was to extract, classify (i.e. label) and cluster segments of the free-text notes that correspond to specific clinical features (e.g., family background, comorbidities or toxicities). Finally, we propose a tool for computer-assisted semantic mapping of segment types to pre-defined ontologies and validate it on a downstream task of category-specific patient similarity. The presented results demonstrate the practical relevance of the proposed approach for building more sophisticated extraction and analytical pipelines deployed on Czech clinical notes.

Petr ZELINA, Vít NOVÁČEK, Jana HALÁMKOVÁ
IEEE BIBM 2023 · DOI ↗ · PDF ↗
TLDR: We test, whether our clinical note segmentation approach works on an English dataset (MIMIC-III)
This paper presents a text-mining approach to extracting and organizing segments from unstructured clinical notes in an unsupervised way. Our work is motivated by the real challenge of poor semantic integration between clinical notes produced by different doctors, departments, or hospitals. This can lead to clinicians overlooking important information, especially for patients with long and varied medical histories. This work extends a previous approach developed for Czech breast cancer patients and validates it on the publicly accessible MIMIC-III English dataset, demonstrating its universal and language-independent applicability. Our work is a stepping stone to a broad array of downstream tasks, such as summarizing or integrating patient records, extracting structured information, or computing patient embeddings. Additionally, the paper presents a clustering analysis of the latent space of note segment types, using hierarchical clustering and an interactive treemap visualization. The presented results demonstrate that this approach generalizes well for MIMIC and English.

Petr ZELINA
Master's thesis · PDF ↗
TLDR: I explore how to use unstructured clinical notes for calculating how similar are different pairs of patients.
This thesis introduces a new way of measuring patient similarity based only on unstructured clinical notes. The system is able to focus on different aspects of patient similarity using a novel semi-supervised note-filtering technique. This approach has been tested on a dataset of Czech clinical notes of 4267 breast cancer patients. A validation study in cooperation with clinicians from the Masaryk Memorial Cancer Institute has been conducted and evaluated. The results show that this system is able to capture some patient similarity categories well, but more research needs to be done to predict others.

Marko Řeháček
Master's thesis · PDF ↗
Despite the growth in the number of breast cancer patients, there is still a lack of research focused on delivering digital solutions to cater their information needs. Patients often lack sufficient information about the disease and tend to feel disempowered. This thesis explores the design and development of a patient-centered eHealth application, which aims to empower breast cancer patients by providing them with comprehensible information about their condition, treatment, medication, and psychocognitive aspects of the care. The research was conducted in close collaboration with a clinical oncologist and the patients. Based on this collaboration, an application was designed and a prototype of the application was tested with the patients.
Adrianna Janik, Maria Torrente, Luca Costabello, Virginia Calvo, Brian Walsh, Carlos Camps, Sameh K Mohamed, Ana L Ortega, Vít Nováček, Bartomeu Massuti, and others
JCO Clinical Cancer Informatics · DOI ↗ · PDF ↗
TLDR: We used machine learning to predict the risk of cancer recurrence in patients with early-stage non-small cell lung cancer. By training models on patient data (both in tables and as graphs), we found that these AI methods can accurately estimate the probability of relapse, which can help doctors personalize treatment and improve patient outcomes.
Stratifying patients with cancer according to risk of relapse can personalize their care. In this work, we provide an answer to the following research question: How to use machine learning to estimate probability of relapse in patients with early-stage non–small-cell lung cancer (NSCLC)? For predicting relapse in 1,387 patients with early-stage (I-II) NSCLC from the Spanish Lung Cancer Group data (average age 65.7 years, female 24.8%, male 75.2%), we train tabular and graph machine learning models. We generate automatic explanations for the predictions of such models. For models trained on tabular data, we adopt SHapley Additive exPlanations local explanations to gauge how each patient feature contributes to the predicted outcome. We explain graph machine learning predictions with an example-based method that highlights influential past patients. Machine learning models trained on tabular data exhibit a 76% accuracy for the random forest model at predicting relapse evaluated with a 10-fold cross-validation. Graph machine learning reaches 68% accuracy over a held-out test set of 200 patients, calibrated on a held-out set of 100 patients. Our results show that machine learning models trained on tabular and graph data can enable objective, personalized, and reproducible prediction of relapse and, therefore, disease outcome in patients with early-stage NSCLC. With further prospective and multisite validation, and additional radiological and molecular data, this prognostic model could potentially serve as a predictive decision support tool for deciding the use of adjuvant treatments in early-stage lung cancer.
Petr ZELINA, Jana HALÁMKOVÁ, Vít NOVÁČEK
IEEE BIBM 2022 · DOI ↗ · PDF ↗
This work is motivated by the scarcity of tools for accurate, unsupervised information extraction from unstructured clinical notes in computationally underrepresented languages, such as Czech. We introduce a stepping stone to a broad array of downstream tasks such as summarisation or integration of individual patient records, extraction of structured information for national cancer registry reporting or building of semi-structured semantic patient representations for computing patient embeddings. More specifically, we present a method for unsupervised extraction of semantically-labelled textual segments from clinical notes and test it out on a dataset of Czech breast cancer patients, provided by Masaryk Memorial Cancer Institute (the largest Czech hospital specialising in oncology). Our goal was to extract, classify (i.e. label) and cluster segments of the free-text notes that correspond to specific clinical features (e.g., family background, comorbidities or toxicities). The presented results demonstrate the practical relevance of the proposed approach for building more sophisticated extraction and analytical pipelines deployed on Czech clinical notes.

Marek TOMA
Master's thesis · PDF ↗
Drug repurposing is a strategy for the development of new treatment options with already existing drugs. Computational approaches based on machine learning can further reduce the cost and accelerate the drug development process. In this work, we develop relational learning models for the prediction of associations between drugs and diseases. First, we design a model trained on the biomedical knowledge graph. Then, we develop two multimodal extensions, that incorporate textual data into the model. Lastly, we compare the performance of our models and analyze the quality of their predictions, showing promising results for some diseases, and we discuss the limitations and their possible solutions.

Sameh K Mohamed, Brian Walsh, Mohan Timilsina, Maria Torrente, Fabio Franco, Mariano Provencio, Adrianna Janik, Luca Costabello, Pasquale Minervini, Pontus Stenetorp, and others
AMIA Annual Symposium Proceedings · DOI ↗ · PDF ↗
TLDR: We used various machine learning models to predict if non-small cell lung cancer will return in patients after early-stage treatment. Our study, using data from over 2400 patients, shows that these models, especially random forest, can effectively predict recurrence, offering a valuable tool for personalized treatment decisions and improving patient outcomes.
Early detection and mitigation of disease recurrence in non-small cell lung cancer (NSCLC) patients is a non-trivial problem. This preliminary study describes an experimental suite of various machine learning models applied to a patient cohort of 2442 early-stage NSCLC patients to predict recurrence. We provide the outcomes of our assessment of basic supervised machine learning methods, evaluating these models on a binary classification task to classify patients with successful treatments into one of two categories. The results show that the random forest model achieved the best results in terms of all the used evaluation metrics. We discuss the promising results achieved, as well as the lessons learned while developing this baseline for further, more advanced studies in this area. Our results show that machine learning models trained on tabular and graph data can enable objective, personalized, and reproducible prediction of relapse and, therefore, disease outcome in patients with early-stage NSCLC. With further prospective and multisite validation, and additional radiological and molecular data, this prognostic model could potentially serve as a predictive decision support tool for deciding the use of adjuvant treatments in early-stage lung cancer.

Maria Torrente, Pedro A Sousa, Roberto Hernácdez, Mariola Blanco, Virginia Calvo, Ana Collazo, Gracinda R Guerreiro, Beatriz Nunez, Joao Pimentao, Juan Cristobal Sánchez, and others
TLDR: We developed an AI tool to better analyze patient data and predict outcomes in cancer. Based on the CLARIFY study, this tool combines various types of patient information to provide more accurate and personalized prognoses, helping doctors make better treatment decisions.
Simple prognostic biomarkers are critical in clinical practice but often overlook the multi-parametric nature of cancer. This study explores the potential of an **Artificial Intelligence (AI)-based tool** to improve data analysis and prognosis in cancer patients, presenting results from the **CLARIFY study**. The CLARIFY study is a real-world, multicenter, observational study collecting extensive clinical, pathological, and molecular data from various cancer types. The AI tool integrates and analyzes this complex, heterogeneous data to identify patterns and generate more accurate prognostic predictions than traditional methods. Specifically, the tool leverages machine learning algorithms to uncover subtle relationships within the high-dimensional data, providing personalized risk stratification and insights into disease progression. Our findings demonstrate that this AI-driven approach significantly enhances prognostic accuracy, potentially leading to more informed treatment decisions and improved patient outcomes by considering the intricate interplay of multiple factors in cancer development and progression.
Vít NOVOTNÝ, Michal ŠTEFÁNIK, Dávid LUPTÁK, Martin GELETKA, Petr ZELINA, Petr SOJKA
CLEF 2021 (CEUR-WS) · PDF ↗
We report on the systems that the Math Information Retrieval group at Masaryk University (MIRMU) and the team of Faculty of Informatics students (MSM) prepared for Task 1 (Find Answers) of the ARQMath lab at the CLEF conference. We have prototyped ten math-aware information retrieval (MIR) systems for the main question-answering task. We ensembled the results of the ten "weak" individual systems into committees and let them vote to provide answers to questions. We evaluated the proposed individual systems and ensembles, considering their diversity, hyperparameters, and representations used, and classified their approaches. We have shown the diversity of all systems and evaluated four voting algorithms to collect and rank the answers. Ensembling techniques consistently outperformed the base systems and showed the power of voting of diverse systems. Our prototypes help to understand the challenging problems of question answering in the STEM domain and our novel reproducible evaluation framework sets a new direction in MIR research. Finally, we formulate ten commandments for future work in the area.
Sameh K Mohamed, Aayah Nounu, Vít Nováček
Briefings in Bioinformatics · DOI ↗ · PDF ↗
TLDR: We review how advanced computer models called 'knowledge graph embeddings' are being used to solve big problems in biology. These models help make sense of vast amounts of biological data to, for example, find new drugs, understand diseases better, or predict how proteins interact, ultimately speeding up scientific discovery.
The rapid growth of biological data, combined with its inherent complexity and heterogeneity, necessitates advanced computational approaches for extracting meaningful insights. **Knowledge graph embedding (KGE) models** have emerged as powerful tools for representing and reasoning over vast biological knowledge. This review provides a comprehensive overview of the diverse **biological applications of knowledge graph embedding models**. We categorize these applications into key areas such as drug discovery and repurposing (e.g., drug-target interaction prediction, adverse drug reaction prediction), disease genomics (e.g., gene-disease association prediction, disease mechanism elucidation), protein-protein interaction prediction, and biological pathway analysis. We discuss the underlying principles of various KGE models, their advantages in handling complex biological relationships, and the challenges associated with their implementation in biological domains. Furthermore, we highlight recent advancements and future directions, emphasizing the transformative potential of KGEs in accelerating biological discovery and precision medicine.
Petr ZELINA
Bachelor's thesis · PDF ↗
TLDR: Experiments with pre-training of a Czech variant of the ALBERT encoder-only transformer.
This thesis explores a new language model called ALBERT, released by Google Research in 2019. The ALBERT architecture has been very successful in English NLP tasks, and this thesis tests its applicability for the Czech language. The work gives an overview of the models leading to ALBERT and their design. Further, the creation process of several Czech ALBERT models is described. Finally, their performance is evaluated, and the models are compared with the current Czech state-of-the-art solutions for text classification and question answering tasks.
Sameh K Mohamed, Vít Nováček, Aayah Nounu
Bioinformatics · DOI ↗ · PDF ↗
TLDR: We developed a new computer method to find which proteins drugs might interact with (drug targets). We do this by building a large network of biomedical facts (a 'knowledge graph') and then using advanced AI to learn patterns within it. Our method can accurately predict new drug-protein interactions, which can help speed up the discovery of new medicines.
Motivation: Identifying protein targets for new or existing drugs is a crucial step in drug discovery and repurposing. Traditional experimental methods are time-consuming and expensive. Computational approaches, particularly those leveraging the growing volume of biomedical data, offer a promising alternative. Results: We propose a novel computational framework that utilizes knowledge graph embeddings (KGEs) for predicting drug–target interactions (DTIs). We construct a comprehensive knowledge graph by integrating diverse biomedical datasets, including information about drugs, proteins, diseases, and pathways. This graph serves as the foundation for learning low-dimensional vector representations (embeddings) of entities and relations using advanced KGE models. By treating DTI prediction as a link prediction task within this knowledge graph, our model can effectively identify novel interactions. We evaluated our method on a benchmark dataset and demonstrated superior performance compared to state-of-the-art computational methods, achieving high accuracy in predicting both known and potential novel drug-protein associations. This framework provides a powerful tool for accelerating drug discovery by efficiently prioritizing candidate drug targets for experimental validation.
Vít Nováček, Gavin McGauran, David Matallanas, Adriáč Vallejo Blanco, Piero Conca, Emir Muñoz, Luca Costabello, Kamalesh Kanakaraj, Zeeshan Nawaz, Brian Walsh, and others
PLoS Computational Biology · DOI ↗ · PDF ↗
TLDR: We developed a new computer method to accurately predict how enzymes called kinases interact with other molecules (substrates). We did this by building a large network of biological information (a 'knowledge graph') and using advanced AI to find hidden patterns. Our method is much better at predicting these interactions, which is important for understanding how cells work and finding new drugs.
Kinases are a crucial class of enzymes involved in almost all cellular processes, making the accurate prediction of kinase-substrate interactions (KSIs) vital for understanding cellular signaling and drug discovery. Traditional methods for identifying KSIs are often time-consuming and labor-intensive. In this study, we present a novel computational approach that leverages **knowledge graphs** to accurately predict kinase-substrate networks. We construct a comprehensive knowledge graph by integrating diverse biological data, including information about kinases, substrates, phosphorylation sites, and relevant biological pathways. By representing these entities and their relationships within a graph, we can apply **knowledge graph embedding (KGE)** techniques to learn latent representations that capture complex biological associations. We then use these embeddings to predict novel KSIs. Our extensive evaluation demonstrates that this knowledge graph-based approach significantly outperforms existing methods in predicting kinase-substrate interactions, providing a powerful tool for accelerating the discovery of new therapeutic targets and advancing our understanding of cellular regulation.
Vít Nováček, Sameh K Mohamed
AMIA Summits on Translational Science Proceedings · DOI ↗ · PDF ↗
TLDR: We developed a computer system to predict harmful side effects when patients take multiple medications at once. By building a network of drug and health information ('knowledge graph') and using advanced AI, our system can identify potential problems, helping doctors prescribe safer drug combinations.
The increasing prevalence of polypharmacy (the concurrent use of multiple medications) in patient care leads to a growing concern about potential adverse drug-drug interactions, often manifesting as polypharmacy side-effects (PSEs). Predicting these complex interactions is a major challenge in clinical practice. This paper presents a novel computational approach to predict polypharmacy side-effects by leveraging **knowledge graph embeddings (KGEs)**. We construct a comprehensive knowledge graph that integrates information about drugs, diseases, symptoms, and known adverse drug reactions. By applying KGE models, we learn low-dimensional vector representations of entities and relationships within this graph. These embeddings are then used to predict novel and previously unobserved PSEs, treating the problem as a link prediction task. Our experiments demonstrate that this knowledge graph-based approach can effectively identify potential polypharmacy side-effects, offering a valuable tool for clinicians to enhance patient safety and optimize medication regimens. The method can help prioritize drug combinations for further investigation and reduce the risk of adverse events.
Brian Walsh, Sameh K Mohamed, Vít Nováček
TLDR: We created **BioKG**, a new, large network of biological information (a 'knowledge graph') that connects different types of data like genes, drugs, and diseases. This graph is designed to help computers learn complex relationships within biological data, which can speed up discoveries, like finding new drug targets or understanding disease-causing genes.
The exponential growth of biological data across various modalities presents both opportunities and challenges for knowledge discovery. Integrating and making sense of this heterogeneous data is crucial for advancing biomedical research. In this paper, we introduce **BioKG**, a novel and comprehensive **knowledge graph specifically designed for relational learning on biological data**. BioKG integrates a wide array of biological entities and relationships from diverse public databases, including information on genes, proteins, diseases, drugs, pathways, and molecular interactions. We detail the schema and construction methodology of BioKG, emphasizing its rich semantic content and interconnectedness. Furthermore, we demonstrate the utility of BioKG by applying state-of-the-art knowledge graph embedding (KGE) models to perform various downstream tasks, such as drug-target interaction prediction and disease gene prioritization. Our results show that BioKG, combined with relational learning techniques, provides a powerful resource for inferring novel biological associations and accelerating discoveries in life sciences.
Emir Muñoz, Vít Nováček, Pierre-Yves Vandenbussche
Briefings in Bioinformatics · DOI ↗ · PDF ↗
TLDR: We created a system that uses a large network of interconnected biomedical facts (a **knowledge graph**) and advanced machine learning to better predict what side effects drugs might have. This helps identify potential problems with drugs more accurately and efficiently.
The increasing volume of biomedical data and the complexity of drug interactions make the prediction of adverse drug reactions (ADRs) a significant challenge. This paper presents a novel approach to facilitate ADR prediction by leveraging **knowledge graphs** and **multi-label learning models**. We construct a comprehensive knowledge graph integrating various types of biomedical information, such as drug properties, protein targets, diseases, and known ADRs. This rich representation allows us to capture intricate relationships between drugs and their potential side effects. Subsequently, we employ multi-label learning techniques, which can predict multiple adverse reactions simultaneously for a given drug, offering a more efficient and accurate prediction framework. Our experimental results demonstrate that this integrated approach significantly improves the accuracy and completeness of ADR prediction compared to traditional methods, providing a valuable tool for pharmacovigilance and drug development.
Sameh K Mohamed, Vít Nováček
TLDR: We improved how computers predict missing connections in knowledge graphs by breaking down the 'meaning' of each item and relationship into several smaller parts. This makes the predictions more accurate, especially for complex relationships.
Link prediction in knowledge graphs aims to infer missing relationships between entities. While existing embedding models have shown promise, they often struggle with complex relations and the ability to capture diverse aspects of relationships. In this paper, we propose a novel approach for link prediction using **multi-part embeddings**. Instead of representing each entity and relation with a single vector, we decompose them into multiple, smaller embedding parts. This allows for a more granular and flexible representation of semantic information, enabling the model to capture different facets of a relationship. Our experiments on benchmark knowledge graphs demonstrate that this multi-part embedding strategy significantly improves the accuracy of link prediction, particularly for relations with varying arities and complex semantic patterns.
Sameh K Mohamed, Aayah Nounu, Vít Nováček
TLDR: We developed a new computer method to find potential drug targets (proteins that drugs interact with). We treat the problem like finding missing connections in a large network of biomedical facts (a 'knowledge graph'). By using advanced machine learning on this network, our method can accurately predict new drug-target relationships, which can speed up drug discovery and research.
Computational methods for predicting drug-target interactions (DTIs) can provide valuable insights into the drug mechanism of action, helping to identify new promising (on-target) or unintended (off-target) effects of drugs. This paper introduces a novel computational approach for predicting drug target proteins by formulating the problem as a link prediction task on knowledge graphs. We process drug and target information as a knowledge graph of interconnected drugs, proteins, diseases, pathways, and other relevant entities. We then apply knowledge graph embedding (KGE) models over this data to enable scoring drug-target associations, employing a customized version of the state-of-the-art KGE model ComplEx. We generate a benchmarking dataset based on the KEGG database to train and evaluate our method. Our experiments show that our method achieves superior results compared to other traditional KGE models, predicting drug target links with a mean reciprocal rank (MRR) of 0.78 and Hits@10 of 0.88. This provides a promising basis for further experimentation and comparisons with domain-specific predictive models.
Sameh K Mohamed, Vít Nováček, Pierre-Yves Vandenbussche, Emir Muñoz
DL4KG@ESWC
TLDR: We examined different mathematical formulas (called **loss functions**) that are used to train computer models for understanding relationships in knowledge graphs. We show how choosing a different loss function affects how well the computer learns and predicts information, helping others pick the best one for their needs.
Knowledge graph embedding (KGE) models learn continuous vector representations of entities and relations within a knowledge graph. A critical component influencing the performance and characteristics of these models is the **loss function** used during training. This paper provides an overview and empirical analysis of various loss functions commonly employed in KGE models, discussing their theoretical underpinnings and practical implications. We examine how different loss functions (e.g., margin-based, logistic, softplus) affect the learning process, the quality of the learned embeddings, and the overall performance on downstream tasks such as link prediction. Through comparative experiments, we highlight the strengths and weaknesses of each loss function, offering insights into their suitability for different types of knowledge graphs and learning objectives. The aim is to guide researchers and practitioners in selecting appropriate loss functions for their specific KGE applications.
Emir Muñoz, Vít Nováček, Pierre-Yves Vandenbussche
AMIA Annual Symposium Proceedings · PDF ↗
TLDR: We investigated how comparing drugs based on their characteristics can help predict unexpected side effects they might cause, improving drug safety and discovery.
The text does not contain a full abstract. The purpose of this work is to explore the use of drug similarities to discover possible adverse reactions, leveraging existing knowledge about drug properties and known side effects to predict new ones. This approach can help in pharmacovigilance and drug development by identifying potential risks earlier.
Pasquale Minervini, Luca Costabello, Emir Muñoz, Vít Nováček, Pierre-Yves Vandenbussche
TLDR: We improved how computers learn about relationships in knowledge graphs (like a network of facts). We did this by adding rules (called 'axioms') that ensure these learned relationships are logically correct, for example, if A is equivalent to B, then B should be equivalent to A. This makes the computer's understanding of the knowledge graph more accurate and reliable.
Knowledge graph embedding models aim to learn low-dimensional representations of entities and relations in knowledge graphs. While these models have achieved impressive empirical results on tasks like link prediction, they often struggle to incorporate logical axioms (e.g., equivalence, inverse properties) that are present in ontologies or knowledge bases. In this paper, we propose a novel approach to regularize knowledge graph embeddings by explicitly enforcing these logical axioms during training. Specifically, we introduce regularization terms that penalize deviations from equivalence and inversion properties, leading to more logically consistent and accurate embeddings. Our experiments demonstrate that this regularization significantly improves the performance of various embedding models on standard benchmarks, especially in scenarios where logical consistency is crucial.
Lenka TEJMLOVÁ, Jiří ŠEBESTA, Petr ZELINA
Radioelektronika 2016 · DOI ↗
Vít Nováček, Siegfried Handschuh, Stefan Decker
TLDR: We propose a new way to understand meaning on the web by adding a 'distributional layer' to existing semantic models. This layer uses statistical patterns of how words appear together to capture their meaning, which helps overcome limitations of traditional semantic approaches like lack of high-quality annotations and incomplete knowledge, ultimately making web semantics more robust.
The scarcity of high-quality annotations in many application scenarios has recently led to an increasing interest in devising learning techniques that combine unlabeled data with labeled data in a network. In this work, we focus on the label propagation problem in multilayer networks. Distributional semantics is defined upon the assumption that the context surrounding a given word in a text provides important information about its meaning. It focuses on the construction of a semantic model for a word based on the statistical distribution of co-located words in texts. These semantic models are naturally represented by vector space models (VSMs), where the meaning of a word can be defined by a weighted vector, which represents the association pattern of co-occurring words in a corpus. This work proposes that distributional semantic models can serve as complementary semantic layers to the relational/logical model, addressing issues of semantic approximation and incompleteness on the web.
Vít Nováček, Tudor Groza, Siegfried Handschuh, Stefan Decker
J. Web Semant. · DOI ↗ · PDF ↗
TLDR: We created **CORAAL**, a system that helps researchers explore scientific papers more effectively. Instead of just searching by keywords, CORAAL uses advanced technology to understand the concepts in papers and how they relate, making it easier to discover connections, understand research areas, and find important information.
The sheer volume of scientific publications makes it challenging for researchers to efficiently discover relevant information and understand the landscape of their field. This paper introduces **CORAAL (COncept-RelAted Access to ALl publications)**, a system designed to facilitate in-depth exploration and knowledge discovery within academic literature. CORAAL goes beyond traditional keyword-based search by leveraging **semantic technologies** to organize and interlink publications based on their underlying concepts and relationships. It allows users to 'dive into publications' by providing concept-centric navigation, semantic Browse, and visualizations that highlight key themes, authors, and connections across articles. This approach helps users to 'bathe in the knowledge' by offering a more intuitive and comprehensive way to grasp complex research areas and identify emerging trends. We discuss the architecture of CORAAL and demonstrate how it enhances information retrieval and knowledge management in the scientific domain.
Vít Nováček, Loredana Laera, Siegfried Handschuh, Brian Davis
Journal of Biomedical Informatics · DOI ↗ · PDF ↗
TLDR: We developed a system to automatically extend biomedical ontologies (structured knowledge bases) by using information extracted from text. This system is designed to handle the constantly changing and large amounts of data in biomedicine, allowing for dynamic updates and integration of knowledge, which is crucial for modern medical research.
This paper introduces a novel ontology integration framework that explicitly takes the dynamics and data-intensiveness of many practical application scenarios into account. The research is motivated by the needs of biomedical scenarios, where the search for semantics-enabled solutions has recently been identified. The authors present a concrete example of the integration process in life sciences settings. The proposed framework considers arguments and propositions specific to the matching task and based on ontology semantics. It deals with proper placement of ontology learning, evaluation and negotiation methods and with integration of learned and collaborative ontologies in a novel way. The transfer possibilities of the framework are justified by elaborated application scenarios from the medicine domain.
Vít Nováček, Loredana Laera, Siegfried Handschuh
IWOD-07 · PDF ↗
TLDR: We developed a system that helps integrate ontologies (structured knowledge bases) learned automatically by computers into a shared environment where people can work together on them. This makes it easier to create and maintain knowledge for the Semantic Web.
Ontologies are commonly considered as one of the essential parts of the Semantic Web vision, providing a theoretical basis and implementation framework for conceptual integration and information sharing among various domains. This work explores the semi-automatic integration of learned ontologies into a collaborative framework. It likely discusses methods and tools that leverage machine learning for ontology creation or refinement, while still allowing for human collaboration and oversight in the integration process. The paper aims to facilitate the evolution and sharing of knowledge in dynamic, collaborative environments.
Vit Nováček, Maciej Dabrowski, Sebastian Ryszard Kruk, Siegfried Handschuh
FLAIRS · PDF ↗
TLDR: We developed a method to automatically suggest additions and improvements to shared community ontologies, aiming to make them more comprehensive and useful without extensive manual effort.
This work explores extending community ontologies by automatically generating suggestions. While the exact abstract wasn't directly found, the title and context suggest research into methods for enriching shared knowledge representations with automated tools. This likely involves analyzing existing data or texts to propose new concepts, relationships, or modifications to an ontology.
Vít Nováček, Pavel Smrž
TLDR: We propose a new way to represent and manage uncertainty within ontologies (structured knowledge bases) that are automatically built. Our framework, called ANUIC, helps improve the accuracy and quality of these learned ontologies, particularly in tasks like creating taxonomies, by accounting for imprecise or incomplete information.
This paper presents new results of our research on uncertainty incorporation into ontologies created automatically by means of Human Language Technologies. The research is related to OLE (Ontology Learning) – a project aimed at bottom-up generation and merging of ontologies. It utilizes a proposal of an expressive fuzzy knowledge representation framework called ANUIC (Adaptive Net of Universally Interrelated Concepts). We discuss our recent achievements in taxonomy acquisition and show how even simple application of the principles of ANUIC can improve the results of initial knowledge extraction methods.