Methodological Issues in Predicting Pediatric Epilepsy Surgery Candidates through Natural Language Processing and Machine Learning

Biomedical Informatics Insights ◽

10.4137/bii.s38308 ◽

2016 ◽

Vol 8 ◽

pp. BII.S38308 ◽

Cited By ~ 17

Author(s):

Kevin Bretonnel Cohen ◽

Benjamin Glass ◽

Hansel M. Greiner ◽

Katherine Holland-Bouley ◽

Shannon Standridge ◽

...

Keyword(s):

Machine Learning ◽

Natural Language Processing ◽

Language Processing ◽

Epilepsy Surgery ◽

Classification Performance ◽

Training Data ◽

Pediatric Epilepsy ◽

Learning Methods ◽

Machine Learning Methods ◽

Pediatric Epilepsy Surgery

Objective: We describe the development and evaluation of a system that uses machine learning and natural language processing techniques to identify potential candidates for surgical intervention for drug-resistant pediatric epilepsy. The data are comprised of free-text clinical notes extracted from the electronic health record (EHR). Both known clinical outcomes from the EHR and manual chart annotations provide gold standards for the patient's status. The following hypotheses are then tested: 1) machine learning methods can identify epilepsy surgery candidates as well as physicians do and 2) machine learning methods can identify candidates earlier than physicians do. These hypotheses are tested by systematically evaluating the effects of the data source, amount of training data, class balance, classification algorithm, and feature set on classifier performance. The results support both hypotheses, with F-measures ranging from 0.71 to 0.82. The feature set, classification algorithm, amount of training data, class balance, and gold standard all significantly affected classification performance. It was further observed that classification performance was better than the highest agreement between two annotators, even at one year before documented surgery referral. The results demonstrate that such machine learning methods can contribute to predicting pediatric epilepsy surgery candidates and reducing lag time to surgery referral.

Download Full-text

Social Media Content Categorization Using Supervised Based Machine Learning Methods and Natural Language Processing in Bangla Language

2020 11th International Conference on Electrical and Computer Engineering (ICECE) ◽

10.1109/icece51571.2020.9393095 ◽

2020 ◽

Author(s):

Md. Rejaul Alam ◽

Afsana Akter ◽

Minhajul Abedin Shafin ◽

Md. Mehedi Hasan ◽

Antara Mahmud

Keyword(s):

Machine Learning ◽

Social Media ◽

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Media Content ◽

Learning Methods ◽

Machine Learning Methods

Download Full-text

A Comparison of Natural Language Processing and Machine Learning Methods for Phishing Email Detection

The 16th International Conference on Availability, Reliability and Security ◽

10.1145/3465481.3469205 ◽

2021 ◽

Author(s):

Panagiotis Bountakas ◽

Konstantinos Koutroumpouchos ◽

Christos Xenakis

Keyword(s):

Machine Learning ◽

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Learning Methods ◽

Machine Learning Methods

Download Full-text

Novel Natural Language Processing and Machine Learning Methods to Characterize Unstructured Patient-Reported Outcomes

SSRN Electronic Journal ◽

10.2139/ssrn.3736178 ◽

2020 ◽

Author(s):

Zhaohua Lu ◽

Jin-ah Sim ◽

Christopher B. Forrest ◽

Kevin R. Krull ◽

Kumar Srivastava ◽

...

Keyword(s):

Machine Learning ◽

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Patient Reported Outcomes ◽

Learning Methods ◽

Machine Learning Methods ◽

Patient Reported

Download Full-text

Tutorial: Machine Learning Methods in Natural Language Processing

Learning Theory and Kernel Machines - Lecture Notes in Computer Science ◽

10.1007/978-3-540-45167-9_47 ◽

2003 ◽

pp. 655-655 ◽

Cited By ~ 3

Author(s):

Michael Collins

Keyword(s):

Machine Learning ◽

Natural Language Processing ◽

Natural Language ◽

Language Processing ◽

Learning Methods ◽

Machine Learning Methods

Download Full-text

Online Review Content Moderation Using Natural Language Processing and Machine Learning Methods : 2021 Systems and Information Engineering Design Symposium (SIEDS)

2021 Systems and Information Engineering Design Symposium (SIEDS) ◽

10.1109/sieds52267.2021.9483739 ◽

2021 ◽

Author(s):

Alicia Doan ◽

Nathan England ◽

Travis Vitello

Keyword(s):

Machine Learning ◽

Natural Language Processing ◽

Natural Language ◽

Engineering Design ◽

Language Processing ◽

Online Review ◽

Learning Methods ◽

Machine Learning Methods ◽

Information Engineering

Download Full-text

Possibility of Autonomous Estimation of Shiba Goat’s Estrus and Non-Estrus Behavior by Machine Learning Methods

Animals ◽

10.3390/ani10050771 ◽

2020 ◽

Vol 10 (5) ◽

pp. 771

Author(s):

Toshiya Arakawa

Keyword(s):

Neural Network ◽

Machine Learning ◽

Random Forest ◽

Markov Models ◽

Tracking System ◽

Video Tracking ◽

Training Data ◽

Support Vector ◽

Learning Methods ◽

Machine Learning Methods

Mammalian behavior is typically monitored by observation. However, direct observation requires a substantial amount of effort and time, if the number of mammals to be observed is sufficiently large or if the observation is conducted for a prolonged period. In this study, machine learning methods as hidden Markov models (HMMs), random forests, support vector machines (SVMs), and neural networks, were applied to detect and estimate whether a goat is in estrus based on the goat’s behavior; thus, the adequacy of the method was verified. Goat’s tracking data was obtained using a video tracking system and used to estimate whether they, which are in “estrus” or “non-estrus”, were in either states: “approaching the male”, or “standing near the male”. Totally, the PC of random forest seems to be the highest. However, The percentage concordance (PC) value besides the goats whose data were used for training data sets is relatively low. It is suggested that random forest tend to over-fit to training data. Besides random forest, the PC of HMMs and SVMs is high. However, considering the calculation time and HMM’s advantage in that it is a time series model, HMM is better method. The PC of neural network is totally low, however, if the more goat’s data were acquired, neural network would be an adequate method for estimation.

Download Full-text

Multi-label pathway prediction based on active dataset subsampling

10.1101/2020.09.14.297424 ◽

2020 ◽

Author(s):

Abdur Rahman M. A. Basher ◽

Steven J. Hallam

Keyword(s):

Machine Learning ◽

Microbial Communities ◽

Performance Metrics ◽

Class Imbalance ◽

Training Data ◽

Great Promise ◽

Biological Organization ◽

Learning Methods ◽

Machine Learning Methods ◽

Pathway Prediction

AbstractMachine learning methods show great promise in predicting metabolic pathways at different levels of biological organization. However, several complications remain that can degrade prediction performance including inadequately labeled training data, missing feature information, and inherent imbalances in the distribution of enzymes and pathways within a dataset. This class imbalance problem is commonly encountered by the machine learning community when the proportion of instances over class labels within a dataset are uneven, resulting in poor predictive performance for underrepresented classes. Here, we present leADS, multi-label learning based on active dataset subsampling, that leverages the idea of subsampling points from a pool of data to reduce the negative impact of training loss due to class imbalance. Specifically, leADS performs an iterative process to: (i)-construct an acquisition model in an ensemble framework; (ii) select informative points using an appropriate acquisition function; and (iii) train on selected samples. Multiple base learners are implemented in parallel where each is assigned a portion of labeled training data to learn pathways. We benchmark leADS using a corpora of 10 experimental datasets manifesting diverse multi-label properties used in previous pathway prediction studies, including manually curated organismal genomes, synthetic microbial communities and low complexity microbial communities. Resulting performance metrics equaled or exceeded previously reported machine learning methods for both organismal and multi-organismal genomes while establishing an extensible framework for navigating class imbalances across diverse real world datasets.Availability and implementationThe software package, and installation instructions are published on github.com/[email protected]

Download Full-text

Machine Learning and Value-Based Software Engineering

Software and Intelligent Sciences ◽

10.4018/978-1-4666-0261-8.ch017 ◽

2012 ◽

pp. 287-301

Author(s):

Du Zhang

Keyword(s):

Machine Learning ◽

Software Engineering ◽

Full Range ◽

Training Data ◽

Learning Methods ◽

Engineering Research ◽

Machine Learning Methods ◽

Machine Learning Applications ◽

Value Propositions ◽

Software Engineering Research

Software engineering research and practice thus far are primarily conducted in a value-neutral setting where each artifact in software development such as requirement, use case, test case, and defect, is treated as equally important during a software system development process. There are a number of shortcomings of such value-neutral software engineering. Value-based software engineering is to integrate value considerations into the full range of existing and emerging software engineering principles and practices. Machine learning has been playing an increasingly important role in helping develop and maintain large and complex software systems. However, machine learning applications to software engineering have been largely confined to the value-neutral software engineering setting. In this paper, the general message to be conveyed is to apply machine learning methods and algorithms to value-based software engineering. The training data or the background knowledge or domain theory or heuristics or bias used by machine learning methods in generating target models or functions should be aligned with stakeholders’ value propositions. An initial research agenda is proposed for machine learning in value-based software engineering.

Download Full-text

Rethinking domain adaptation for machine learning over clinical language

JAMIA Open ◽

10.1093/jamiaopen/ooaa010 ◽

2020 ◽

Vol 3 (2) ◽

pp. 146-150

Author(s):

Egoitz Laparra ◽

Steven Bethard ◽

Timothy A Miller

Keyword(s):

Machine Learning ◽

Natural Language Processing ◽

Language Processing ◽

Domain Adaptation ◽

Positive Impact ◽

Training Data ◽

Clinical Use ◽

Research Directions ◽

Clinical Natural Language Processing ◽

Support Research

Abstract Building clinical natural language processing (NLP) systems that work on widely varying data is an absolute necessity because of the expense of obtaining new training data. While domain adaptation research can have a positive impact on this problem, the most widely studied paradigms do not take into account the realities of clinical data sharing. To address this issue, we lay out a taxonomy of domain adaptation, parameterizing by what data is shareable. We show that the most realistic settings for clinical use cases are seriously under-studied. To support research in these important directions, we make a series of recommendations, not just for domain adaptation but for clinical NLP in general, that ensure that data, shared tasks, and released models are broadly useful, and that initiate research directions where the clinical NLP community can lead the broader NLP and machine learning fields.

Download Full-text

TADA: phylogenetic augmentation of microbiome samples enhances phenotype classification

Bioinformatics ◽

10.1093/bioinformatics/btz394 ◽

2019 ◽

Vol 35 (14) ◽

pp. i31-i40 ◽

Cited By ~ 1

Author(s):

Erfan Sayyari ◽

Ban Kawas ◽

Siavash Mirarab

Keyword(s):

Machine Learning ◽

Sample Size ◽

Data Augmentation ◽

Training Data ◽

Supplementary Information ◽

High Dimensional ◽

Learning Methods ◽

Machine Learning Methods ◽

Phenotype Classification ◽

Microbiome Data

Abstract Motivation Learning associations of traits with the microbial composition of a set of samples is a fundamental goal in microbiome studies. Recently, machine learning methods have been explored for this goal, with some promise. However, in comparison to other fields, microbiome data are high-dimensional and not abundant; leading to a high-dimensional low-sample-size under-determined system. Moreover, microbiome data are often unbalanced and biased. Given such training data, machine learning methods often fail to perform a classification task with sufficient accuracy. Lack of signal is especially problematic when classes are represented in an unbalanced way in the training data; with some classes under-represented. The presence of inter-correlations among subsets of observations further compounds these issues. As a result, machine learning methods have had only limited success in predicting many traits from microbiome. Data augmentation consists of building synthetic samples and adding them to the training data and is a technique that has proved helpful for many machine learning tasks. Results In this paper, we propose a new data augmentation technique for classifying phenotypes based on the microbiome. Our algorithm, called TADA, uses available data and a statistical generative model to create new samples augmenting existing ones, addressing issues of low-sample-size. In generating new samples, TADA takes into account phylogenetic relationships between microbial species. On two real datasets, we show that adding these synthetic samples to the training set improves the accuracy of downstream classification, especially when the training data have an unbalanced representation of classes. Availability and implementation TADA is available at https://github.com/tada-alg/TADA. Supplementary information Supplementary data are available at Bioinformatics online.

Download Full-text