Hashing Hyperplane Queries to Near Points with Applications to Large-Scale Active Learning

2014 ◽  
Vol 36 (2) ◽  
pp. 276-288 ◽  
Author(s):  
Sudheendra Vijayanarasimhan ◽  
Prateek Jain ◽  
Kristen Grauman
Keyword(s):  
2019 ◽  
Author(s):  
Kyle Konze ◽  
Pieter Bos ◽  
Markus Dahlgren ◽  
Karl Leswing ◽  
Ivan Tubert-Brohman ◽  
...  

We report a new computational technique, PathFinder, that uses retrosynthetic analysis followed by combinatorial synthesis to generate novel compounds in synthetically accessible chemical space. Coupling PathFinder with active learning and cloud-based free energy calculations allows for large-scale potency predictions of compounds on a timescale that impacts drug discovery. The process is further accelerated by using a combination of population-based statistics and active learning techniques. Using this approach, we rapidly optimized R-groups and core hops for inhibitors of cyclin-dependent kinase 2. We explored greater than 300 thousand ideas and identified 35 ligands with diverse commercially available R-groups and a predicted IC<sub>50</sub> < 100 nM, and four unique cores with a predicted IC<sub>50</sub> < 100 nM. The rapid turnaround time, and scale of chemical exploration, suggests that this is a useful approach to accelerate the discovery of novel chemical matter in drug discovery campaigns.


2019 ◽  
Author(s):  
Kyle Konze ◽  
Pieter Bos ◽  
Markus Dahlgren ◽  
Karl Leswing ◽  
Ivan Tubert-Brohman ◽  
...  

We report a new computational technique, PathFinder, that uses retrosynthetic analysis followed by combinatorial synthesis to generate novel compounds in synthetically accessible chemical space. Coupling PathFinder with active learning and cloud-based free energy calculations allows for large-scale potency predictions of compounds on a timescale that impacts drug discovery. The process is further accelerated by using a combination of population-based statistics and active learning techniques. Using this approach, we rapidly optimized R-groups and core hops for inhibitors of cyclin-dependent kinase 2. We explored greater than 300 thousand ideas and identified 35 ligands with diverse commercially available R-groups and a predicted IC<sub>50</sub> < 100 nM, and four unique cores with a predicted IC<sub>50</sub> < 100 nM. The rapid turnaround time, and scale of chemical exploration, suggests that this is a useful approach to accelerate the discovery of novel chemical matter in drug discovery campaigns.


2021 ◽  
Author(s):  
Kai Xu ◽  
Lei Yan ◽  
Bingran You

Force field is a central requirement in molecular dynamics (MD) simulation for accurate description of the potential energy landscape and the time evolution of individual atomic motions. Most energy models are limited by a fundamental tradeoff between accuracy and speed. Although ab initio MD based on density functional theory (DFT) has high accuracy, its high computational cost prevents its use for large-scale and long-timescale simulations. Here, we use Bayesian active learning to construct a Gaussian process model of interatomic forces to describe Pt deposited on Ag(111). An accurate model is obtained within one day of wall time after selecting only 126 atomic environments based on two- and three-body interactions, providing mean absolute errors of 52 and 142 meV/Å for Ag and Pt, respectively. Our work highlights automated and minimalistic training of machine-learning force fields with high fidelity to DFT, which would enable large-scale and long-timescale simulations of alloy surfaces at first-principles accuracy.


Author(s):  
Shaolei Wang ◽  
Zhongyuan Wang ◽  
Wanxiang Che ◽  
Sendong Zhao ◽  
Ting Liu

Spoken language is fundamentally different from the written language in that it contains frequent disfluencies or parts of an utterance that are corrected by the speaker. Disfluency detection (removing these disfluencies) is desirable to clean the input for use in downstream NLP tasks. Most existing approaches to disfluency detection heavily rely on human-annotated data, which is scarce and expensive to obtain in practice. To tackle the training data bottleneck, in this work, we investigate methods for combining self-supervised learning and active learning for disfluency detection. First, we construct large-scale pseudo training data by randomly adding or deleting words from unlabeled data and propose two self-supervised pre-training tasks: (i) a tagging task to detect the added noisy words and (ii) sentence classification to distinguish original sentences from grammatically incorrect sentences. We then combine these two tasks to jointly pre-train a neural network. The pre-trained neural network is then fine-tuned using human-annotated disfluency detection training data. The self-supervised learning method can capture task-special knowledge for disfluency detection and achieve better performance when fine-tuning on a small annotated dataset compared to other supervised methods. However, limited in that the pseudo training data are generated based on simple heuristics and cannot fully cover all the disfluency patterns, there is still a performance gap compared to the supervised models trained on the full training dataset. We further explore how to bridge the performance gap by integrating active learning during the fine-tuning process. Active learning strives to reduce annotation costs by choosing the most critical examples to label and can address the weakness of self-supervised learning with a small annotated dataset. We show that by combining self-supervised learning with active learning, our model is able to match state-of-the-art performance with just about 10% of the original training data on both the commonly used English Switchboard test set and a set of in-house annotated Chinese data.


Sensors ◽  
2020 ◽  
Vol 20 (17) ◽  
pp. 4975
Author(s):  
Fangyu Shi ◽  
Zhaodi Wang ◽  
Menghan Hu ◽  
Guangtao Zhai

Relying on large scale labeled datasets, deep learning has achieved good performance in image classification tasks. In agricultural and biological engineering, image annotation is time-consuming and expensive. It also requires annotators to have technical skills in specific areas. Obtaining the ground truth is difficult because natural images are expensive. In addition, images in these areas are usually stored as multichannel images, such as computed tomography (CT) images, magnetic resonance images (MRI), and hyperspectral images (HSI). In this paper, we present a framework using active learning and deep learning for multichannel image classification. We use three active learning algorithms, including least confidence, margin sampling, and entropy, as the selection criteria. Based on this framework, we further introduce an “image pool” to make full advantage of images generated by data augmentation. To prove the availability of the proposed framework, we present a case study on agricultural hyperspectral image classification. The results show that the proposed framework achieves better performance compared with the deep learning model. Manual annotation of all the training sets achieves an encouraging accuracy. In comparison, using active learning algorithm of entropy and image pool achieves a similar accuracy with only part of the whole training set manually annotated. In practical application, the proposed framework can remarkably reduce labeling effort during the model development and upadting processes, and can be applied to multichannel image classification in agricultural and biological engineering.


2014 ◽  
Vol 11 (1) ◽  
pp. 259-263 ◽  
Author(s):  
Naif Alajlan ◽  
Edoardo Pasolli ◽  
Farid Melgani ◽  
Andrea Franzoso

Author(s):  
Elina Penttinen ◽  
Marjut Jyrkinen

This study aims to examine the suitability of feminist student-centred active learning pedagogy in large-scale classroom settings in a contemporary neoliberalist university context. In the current individualist culture in the academia where students implicitly have adopted a customer-like mind-set, they need to be rational in terms of what they study and how they use their time. Weargue that feminist values are what makes student-centred active learning successful and will enhance the academic expertise of students. However, the values of inclusiveness, low-hierarchy, co-construction of knowledge, and empowerment of feminist pedagogy need to be revisited in the contemporary context. Low-hierarchy may signal to students that they have the ‘upper’hand. Instead of engaging actively in the classroom, they challenge the course content and pedagogical practices. On the basis of our case study data, we claim that this attitude is inherently gendered. Thus, paradoxically, teachers in feminist classrooms need to be careful about the role of ‘service provider’ and assume more assertive leadership roles in order to ensure successful learning outcomes.


Sign in / Sign up

Export Citation Format

Share Document