Computer-implemented methods and systems train, dynamically, a machine-learning network from a base system. Computing learned parameters for the network comprises, for at least a first portion of the machine-learning network, a backpropagation pass through the machine-learning network. The back-propagation pass comprises, for the first portion of the machine-learning network, computation of derivatives, with respect to a loss function, for the learned parameters. The method further comprises making a sensibility level assessment that comprises a determination of whether the machine-learning network produces an insensible result according to a criterion of sensibility. The method further comprises making one or more sensibility-improving modifications in response to a determination, in the sensibility level assessment of the machine-learning network, that the machine-learning network produces an insensible result, such that the one or one or more sensibility-improving modifications make the machine-learning network less vulnerable to producing insensible results.
A diverse set of neural networks are trained to be individually robust against adversarial attacks and diverse in a manner that decreases the ability of an adversarial example to fool the full diverse set. The systems/methods use a diversity criterion that is specialized for measuring diversity in response to adversarial attacks rather than diversity in the classification results. Also, one or more networks can be trained that are less robust to adversarial attacks to use as a diagnostic to detect the presence of an adversarial attack. Also, node-to-node relation regularization links can be used to train diverse networks that are randomly selected from a family of diverse networks with exponentially many members.
Computer-implemented systems and method train a generator and a discriminator, through machine learning, where the generator and discriminator are trained in an adversarial relationship using a simulated, multi-player game. The model parameters for the generator and the discriminator can be updated non-simultaneously. Also, the simulated, multi-player game may comprise a two-person, zero-sum game.
Multi-stage hybrid network integrates relationship regularization links and explainable elements to improve alignment with human values, explainability, robustness, and efficiency. The network comprises neural components, event prediction elements, and probability models across multiple stages, with relationship constraints enforcing structured knowledge representation. Explainable elements provide interpretable rationales for decisions, enhancing transparency. Training incorporates supervised learning, human-guided refinement, semi-automated knowledge engineering, and adversarial robustness techniques. A Socratic reasoning module detects contradictions and refines outputs for logical consistency. Indexed model elements enable dynamic memory optimization for improved efficiency. Candidate outputs may be scored, verified, or selected using neural and symbolic criteria. The invention supports retry loops and configurable subsystem pipelines to improve output quality. Applications include text generation, speech recognition, translation, and decision support. By combining structured constraints, human oversight, and modular architectures, the system improves the trustworthiness, safety, and adaptability of AI systems across diverse modalities and tasks.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
G06N 3/082 - Méthodes d'apprentissage modifiant l’architecture, p. ex. par ajout, suppression ou mise sous silence de nœuds ou de connexions
G06F 16/901 - IndexationStructures de données à cet effetStructures de stockage
G06F 18/21 - Conception ou mise en place de systèmes ou de techniquesExtraction de caractéristiques dans l'espace des caractéristiquesSéparation aveugle de sources
G06F 18/214 - Génération de motifs d'entraînementProcédés de Bootstrapping, p. ex. ”bagging” ou ”boosting”
G06N 20/20 - Techniques d’ensemble en apprentissage automatique
H04L 67/142 - Gestion des états de session pour les protocoles sans étatÉtats des sessions de signalisationSignalisation des états de sessionMécanismes de conservation d’état
6.
SYSTEM AND METHOD FOR IMPROVING MACHINE LEARNING CLASSIFIERS USING SYNTHETIC INPUTS AND GRADIENT DIRECTION ANALYSIS
Computer-implemented systems and methods improve training of a neural network. Whether a target node is not decisive on a training data item is determined. Upon a determination that the target node is not decisive, a partial derivative of an objective for the target node is multiplied by a factor greater than 1.0 for the training data item. Determining whether the target node is not decisive can comprise determining whether a direction of the derivative is in a direction that would cause an update of learned parameters for the network to increase the difference between the activation value of the first target node for the training data item and a neutral activation value for the target node.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06N 3/084 - Rétropropagation, p. ex. suivant l’algorithme du gradient
G06N 7/01 - Modèles graphiques probabilistes, p. ex. réseaux probabilistes
Data-dependent node-to-node knowledge sharing to increase the interpretability of the activation pattern of one or more nodes in a neural network, is implemented by a set of knowledge sharing links. Each link may comprise a knowledge providing node or other source P and a knowledge receiving node R. A knowledge sharing link can impose a node-specific regularization on the knowledge receiving node R to help guide the knowledge receiving node R to have an activation pattern that is more easily interpreted. The specification and training of the knowledge sharing links may be controlled by a cooperative human-AI learning supervisor system in which a human and an artificial intelligence system work cooperatively to improve the interpretability and performance of the client system.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06F 18/40 - Dispositions logicielles spécialement adaptées à la reconnaissance des formes, p. ex. interfaces utilisateur ou boîtes à outils à cet effet
8.
Adaptively training of neural networks via an intelligent learning management system
(a) computing for each datum in a set of training data, activation values for nodes in the neural network and estimates of partial derivatives of an objective function for the neural network for the nodes in the neural network; (b) selecting a target node of the neural network and/or a target datum in the set of training data; (c) selecting a target-specific improvement model for the neural network, wherein the target-specific improvement model, when added to the neural network, improves performance of the neural network for the target node and/or the target datum, as the case may be; (d) training the target-specific improvement model; (e) merging the target-specific improvement model with the neural network to form an expanded neural network; and (f) training the expanded neural network.
Computer systems and methods train a deep neural network through machine learning. In response to detection of a training condition, computer system replaces a target node of the network with a split detector compound node, where, prior to replacement, the target node detected a pattern that activated the target node beyond a specified threshold. The split detector compound node comprises first and second nodes, such that: the first node is activated when significant evidence exists in favor of detection of the pattern in inputs to the first node; and the second node is activated when significant evidence exists against detection of the pattern in inputs to the second node, such that activations of the first and second nodes are computed independently. After replacing the target node with the split detector compound node, training of the network through machine learning is resumed.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
G06N 3/082 - Méthodes d'apprentissage modifiant l’architecture, p. ex. par ajout, suppression ou mise sous silence de nœuds ou de connexions
G06F 16/901 - IndexationStructures de données à cet effetStructures de stockage
G06F 18/21 - Conception ou mise en place de systèmes ou de techniquesExtraction de caractéristiques dans l'espace des caractéristiquesSéparation aveugle de sources
G06F 18/214 - Génération de motifs d'entraînementProcédés de Bootstrapping, p. ex. ”bagging” ou ”boosting”
G06N 20/20 - Techniques d’ensemble en apprentissage automatique
H04L 67/142 - Gestion des états de session pour les protocoles sans étatÉtats des sessions de signalisationSignalisation des états de sessionMécanismes de conservation d’état
Computer-implemented systems and methods improve training of a neural network. Whether a target node is not decisive on a training data item is determined. Upon a determination that the target node is not decisive, a partial derivative of an objective for the target node is multiplied by a factor greater than 1.0 for the training data item. Determining whether the target node is not decisive can comprise determining whether a direction of the derivative is in a direction that would cause an update of learned parameters for the network to increase the difference between the activation value of the first target node for the training data item and a neutral activation value for the target node.
G06N 3/047 - Réseaux probabilistes ou stochastiques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06N 3/084 - Rétropropagation, p. ex. suivant l’algorithme du gradient
G06N 3/088 - Apprentissage non supervisé, p. ex. apprentissage compétitif
G06N 7/01 - Modèles graphiques probabilistes, p. ex. réseaux probabilistes
G06F 12/0815 - Protocoles de cohérence de mémoire cache
G06F 17/18 - Opérations mathématiques complexes pour l'évaluation de données statistiques
Computer systems and computer-implemented methods modify a machine learning network, such as a deep neural network, to introduce judgment to the network. A “combining” node is added to the network, to thereby generate a modified network, where activation of the combining node is based, at least in part, on output from a subject node of the network. The computer system then trains the modified network by, for each training data item in a set of training data, performing forward and back propagation computations through the modified network, where the backward propagation computation through the modified network comprises computing estimated partial derivatives of an error function of an objective for the network, except that the combining node selectively blocks back-propagation of estimated partial derivatives to the subject node, even though activation of the combining node is based on the activation of the subject node.
(a) computing for each datum in a set of training data, activation values for nodes in the neural network and estimates of partial derivatives of an objective function for the neural network for the nodes in the neural network; (b) selecting a target node of the neural network and/or a target datum in the set of training data; (c) selecting a target-specific improvement model for the neural network, wherein the target-specific improvement model, when added to the neural network, improves performance of the neural network for the target node and/or the target datum, as the case may be; (d) training the target-specific improvement model; (e) merging the target-specific improvement model with the neural network to form an expanded neural network; and (f) training the expanded neural network.
Computer-implemented systems and method train a generator and a discriminator, through machine learning, where the generator and discriminator are trained in an adversarial relationship using a simulated, multi-player game. The model parameters for the generator and the discriminator can be updated non-simultaneously. Also, the simulated, multi-player game may comprise a two-person, zero-sum game.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
G06N 3/082 - Méthodes d'apprentissage modifiant l’architecture, p. ex. par ajout, suppression ou mise sous silence de nœuds ou de connexions
G06F 16/901 - IndexationStructures de données à cet effetStructures de stockage
G06F 18/21 - Conception ou mise en place de systèmes ou de techniquesExtraction de caractéristiques dans l'espace des caractéristiquesSéparation aveugle de sources
G06F 18/214 - Génération de motifs d'entraînementProcédés de Bootstrapping, p. ex. ”bagging” ou ”boosting”
G06N 20/20 - Techniques d’ensemble en apprentissage automatique
H04L 67/142 - Gestion des états de session pour les protocoles sans étatÉtats des sessions de signalisationSignalisation des états de sessionMécanismes de conservation d’état
Computer systems and methods train a deep neural network through machine learning. In response to detection of a training condition, computer system replaces a target node of the network with a split detector compound node, where, prior to replacement, the target node detected a pattern that activated the target node beyond a specified threshold. The split detector compound node comprises first and second nodes, such that: the first node is activated when significant evidence exists in favor of detection of the pattern in inputs to the first node; and the second node is activated when significant evidence exists against detection of the pattern in inputs to the second node, such that activations of the first and second nodes are computed independently. After replacing the target node with the split detector compound node, training of the network through machine learning is resumed.
Computer-implemented methods and systems make a generative AI system more explainable. A programmed computer system grows a generative AI system by adding one or more explainable network elements to the generative AI system. Each explainable network element can be trained to discriminate two or more explainable sets of training data items for the generative AI system. After adding the one or more explainable network elements, training of the generative AI system can be updated with the one or more explainable network elements added. Then the programmed computer system can determined whether continued growth of the generative AI system is required.
Data-dependent node-to-node knowledge sharing to increase the interpretability of the activation pattern of one or more nodes in a neural network, is implemented by a set of knowledge sharing links. Each link may comprise a knowledge providing node or other source P and a knowledge receiving node R. A knowledge sharing link can impose a node-specific regularization on the knowledge receiving node R to help guide the knowledge receiving node R to have an activation pattern that is more easily interpreted. The specification and training of the knowledge sharing links may be controlled by a cooperative human-AI learning supervisor system in which a human and an artificial intelligence system work cooperatively to improve the interpretability and performance of the client system.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06F 18/40 - Dispositions logicielles spécialement adaptées à la reconnaissance des formes, p. ex. interfaces utilisateur ou boîtes à outils à cet effet
Computer-implemented methods and systems train, dynamically, a machine-learning network from a base system. Computing learned parameters for the network comprises, for at least a first portion of the machine-learning network, a backpropagation pass through the machine-learning network. The back-propagation pass comprises, for the first portion of the machine-learning network, computation of derivatives, with respect to a loss function, for the learned parameters. The method further comprises making a sensibility level assessment that comprises a determination of whether the machine-learning network produces an insensible result according to a criterion of sensibility. The method further comprises making one or more sensibility-improving modifications in response to a determination, in the sensibility level assessment of the machine-learning network, that the machine-learning network produces an insensible result, such that the one or one or more sensibility-improving modifications make the machine-learning network less vulnerable to producing insensible results.
G06N 3/006 - Vie artificielle, c.-à-d. agencements informatiques simulant la vie fondés sur des formes de vie individuelles ou collectives simulées et virtuelles, p. ex. simulations sociales ou optimisation par essaims particulaires [PSO]
20.
Data-dependent node-to-node knowledge sharing by regularization in deep learning
Data-dependent node-to-node knowledge sharing to increase the interpretability of the activation pattern of one or more nodes in a neural network, is implemented by a set of knowledge sharing links. Each link may comprise a knowledge providing node or other source P and a knowledge receiving node R. A knowledge sharing link can impose a node-specific regularization on the knowledge receiving node R to help guide the knowledge receiving node R to have an activation pattern that is more easily interpreted. The specification and training of the knowledge sharing links may be controlled by a cooperative human-AI learning supervisor system in which a human and an artificial intelligence system work cooperatively to improve the interpretability and performance of the client system.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06F 18/40 - Dispositions logicielles spécialement adaptées à la reconnaissance des formes, p. ex. interfaces utilisateur ou boîtes à outils à cet effet
A machine learning (ML) system includes a student ML system, a learning coach ML system, and a reference system that generates training data for the student ML system. The learning coach ML system learns to make an enhancement to the student ML system or to its learning process, such as updated hyperparameter or a network structural change, based on training of the student ML system with the training data generated by the reference system. The system may also comprise a learning experimentation system that communicates with the reference system to conduct experiments on the learning of the student learning system. Also, the learning experimentation system can determine a cost function for the learning coach ML system.
Computer-implemented methods and systems train, dynamically, a machine-learning network from a base system. Computing learned parameters for the network comprises, for at least a first portion of the machine-learning network, a back-propagation pass through the machine-learning network. The back-propagation pass comprises, for the first portion of the machine-learning network, computation of derivatives, with respect to a loss function, for the learned parameters. The method further comprises making a sensibility level assessment that comprises a determination of whether the machine-learning network produces an insensible result according to a criteria of sensibility. The method further comprises making one or more sensibility-improving modifications in response to a determination, in the sensibility level assessment of the machine-learning network, that the machine-learning network produces an insensible result, such that the one or one or more sensibility-improving modifications make the machine-learning network less vulnerable to producing insensible results.
Computer systems and computer-implemented methods train a neural network iteratively training, through machine learning. The iterative training comprises imposing a first is-not-equal-to regularization link between first and second nodes, where imposing the first is-not-equal-to regularization link between the two nodes comprises adding, during back-propagation of partial derivatives through the neural network for a datum in a training data set, a first regularization cost to a network error loss function for the first node that is inversely proportional to a difference between an activation value for the first node for the datum and an activation value for the second node for the datum.
Computer systems and computer-implemented methods for modifying a machine learning network, such as a deep neural network, to introduce judgment to the network are disclosed. A “combining” node is added to the network, to thereby generate a modified network, where activation of the combining node is based, at least in part, on output from a subject node of the network. The computer system then trains the modified network by, for each training data item in a set of training data, performing forward and back propagation computations through the modified network, where the backward propagation computation through the modified network comprises computing estimated partial derivatives of an error function of an objective for the network, except that the combining node selectively blocks back-propagation of estimated partial derivatives to the subject node, even though activation of the combining node is based on the activation of the subject node.
Computer systems and methods train a deep neural network through machine learning. In response to detection of a training condition, computer system replaces a target node of the network with a split detector compound node, where, prior to replacement, the target node detected a pattern that activated the target node beyond a specified threshold. The split detector compound node comprises first and second nodes, such that: the first node is activated when significant evidence exists in favor of detection of the pattern in inputs to the first node; and the second node is activated when significant evidence exists against detection of the pattern in inputs to the second node, such that activations of the first and second nodes are computed independently. After replacing the target node with the split detector compound node, training of the network through machine learning is resumed.
Computer systems and computer-implemented methods improve a base neural network. In an initial training, preliminary activations values computed for base network nodes for data in the training data set are stored in memory. After the initial training, a new node set is merged into the base neural network to form an expanded neural network, including directly connecting each of the nodes of the new node set to one or more base network nodes. Then the expanded neural network is trained on the training data set using a network error loss function for the expanded neural network. Training the expanded neural network comprises imposing a node-to-node relationship regularization for at least one base network node in the expanded neural network, where imposing the node-to-node relationship regularization comprises adding, during back-propagation of partial derivatives through the expanded neural network for a datum in the training data set, a regularization cost to the network error loss function for the at least one base network node based on a specified relationship between a stored preliminary activation value for the base network node for the datum and an activation value for the base network node of the expanded neural network for the datum.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
G06N 3/082 - Méthodes d'apprentissage modifiant l’architecture, p. ex. par ajout, suppression ou mise sous silence de nœuds ou de connexions
G06F 16/901 - IndexationStructures de données à cet effetStructures de stockage
G06F 18/21 - Conception ou mise en place de systèmes ou de techniquesExtraction de caractéristiques dans l'espace des caractéristiquesSéparation aveugle de sources
G06F 18/214 - Génération de motifs d'entraînementProcédés de Bootstrapping, p. ex. ”bagging” ou ”boosting”
G06N 20/20 - Techniques d’ensemble en apprentissage automatique
H04L 67/142 - Gestion des états de session pour les protocoles sans étatÉtats des sessions de signalisationSignalisation des états de sessionMécanismes de conservation d’état
28.
Data-dependent node-to-node knowledge sharing by regularization in deep learning
Data-dependent node-to-node knowledge sharing to increase the interpretability of the activation pattern of one or more nodes in a neural network, is implemented by a set of knowledge sharing links. Each link may comprise a knowledge providing node or other source P and a knowledge receiving node R. A knowledge sharing link can impose a node-specific regularization on the knowledge receiving node R to help guide the knowledge receiving node R to have an activation pattern that is more easily interpreted. The specification and training of the knowledge sharing links may be controlled by a cooperative human-AI learning supervisor system in which a human and an artificial intelligence system work cooperatively to improve the interpretability and performance of the client system.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06F 18/40 - Dispositions logicielles spécialement adaptées à la reconnaissance des formes, p. ex. interfaces utilisateur ou boîtes à outils à cet effet
Computer systems and computer-implemented methods modify a machine learning network, such as a deep neural network, to introduce judgment to the network. A “combining” node is added to the network, to thereby generate a modified network, where activation of the combining node is based, at least in part, on output from a subject node of the network. The computer system then trains the modified network by, for each training data item in a set of training data, performing forward and back propagation computations through the modified network, where the backward propagation computation through the modified network comprises computing estimated partial derivatives of an error function of an objective for the network, except that the combining node selectively blocks back-propagation of estimated partial derivatives to the subject node, even though activation of the combining node is based on the activation of the subject node.
Computer-implemented systems and method train a generator and a discriminator, through machine learning, where the generator and discriminator are trained in an adversarial relationship using a simulated, multi-player game. The model parameters for the generator and the discriminator can be updated non-simultaneously. Also, the simulated, multi-player game may comprise a two-person, zero-sum game.
A diverse set of neural networks are trained to be individually robust against adversarial attacks and diverse in a manner that decreases the ability of an adversarial example to fool the full diverse set. The systems/methods use a diversity criterion that is specialized for measuring diversity in response to adversarial attacks rather than diversity in the classification results. Also, one or more networks can be trained that are less robust to adversarial attacks to use as a diagnostic to detect the presence of an adversarial attack. Also, node-to-node relation regularization links can be used to train diverse networks that are randomly selected from a family of diverse networks with exponentially many members.
Systems and methods improve performance of a classifier, which comprises a neural network and is trained through machine learning. First and second scores are computed, by the classifier, for each a multiple data examples from a generator. The first score is indicative of whether the data example belongs to a first data cluster and the second score is indicative of whether the data example belongs to a second data cluster. The generator is trained with an objective such that, for each data example generated by the generator, the first and second scores computed by the classifier are equal. Partial derivatives from the classifier are back-propagated for multiple data examples generated by the generator, to obtain a vector, for each data example, that is orthogonal to a decision surface for the classifier. A problem with the classifier is detected based on changes in directions of the vectors. Upon detecting a problem, the classifier is adjusted to reduce errors by the classifier caused by overfitting training data.
G06N 3/047 - Réseaux probabilistes ou stochastiques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06N 3/084 - Rétropropagation, p. ex. suivant l’algorithme du gradient
G06N 7/01 - Modèles graphiques probabilistes, p. ex. réseaux probabilistes
Computer systems and computer-implemented methods modify a machine learning network, such as a deep neural network, to introduce judgment to the network. A “combining” node is added to the network, to thereby generate a modified network, where activation of the combining node is based, at least in part, on output from a subject node of the network. The computer system then trains the modified network by, for each training data item in a set of training data, performing forward and back propagation computations through the modified network, where the backward propagation computation through the modified network comprises computing estimated partial derivatives of an error function of an objective for the network, except that the combining node selectively blocks back-propagation of estimated partial derivatives to the subject node, even though activation of the combining node is based on the activation of the subject node.
A machine learning system includes a coach machine learning system that uses machine learning to help a student machine learning system learn its system. By monitoring the student learning system, the coach machine learning system can learn (through machine learning techniques) “hyperparameters” for the student learning system that control the machine learning process for the student learning system. The machine learning coach could also determine structural modifications for the student learning system architecture. The learning coach can also control data flow to the student learning system.
Computer systems and methods modify a base deep neural network (DNN). The method comprises replacing the target node of the base DNN with a compound node to thereby create a modified base DNN. The compound node comprises at least first and second nodes. The first node is trained to detect target node patterns in inputs to the first node and the second node is trained to detect an absence of the target node patterns in inputs to the second node, and the first and second nodes are trained to be non-complementary. Replacing the target node with the compound node comprises: connecting the first node to the upper sub-network of the base DNN, such that the first node has a weighted connection for each of the one or more base connection that the target node had to the upper sub-network in the base DNN; and connecting the second node to the upper sub-network of the base deep DNN, such that the second node has a weighted connection for each of the one or more base connection that the target node had to the upper sub-network in the base deep DNN. After replacing the target node, the modified base DNN is trained.
Data-dependent node-to-node knowledge sharing to increase the interpretability of the activation pattern of one or more nodes in a neural network, is implemented by a set of knowledge sharing links. Each link may comprise a knowledge providing node or other source P and a knowledge receiving node R. A knowledge sharing link can impose a node-specific regularization on the knowledge receiving node R to help guide the knowledge receiving node R to have an activation pattern that is more easily interpreted. The specification and training of the knowledge sharing links may be controlled by a cooperative human-AI learning supervisor system in which a human and an artificial intelligence system work cooperatively to improve the interpretability and performance of the client system.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06F 18/40 - Dispositions logicielles spécialement adaptées à la reconnaissance des formes, p. ex. interfaces utilisateur ou boîtes à outils à cet effet
A machine learning system includes a coach machine learning system that uses machine learning to help a student machine learning system learn its system. By monitoring the student learning system, the coach machine learning system can learn (through machine learning techniques) “hyperparameters” for the student learning system that control the machine learning process for the student learning system. The machine learning coach could also determine structural modifications for the student learning system architecture. The learning coach can also control data flow to the student learning system.
Machine-learning computer system breaks a neural network into a plurality of modules and tracks the training process module-by-module and datum-by-datum, recording auxiliary information during one iteration of the training process for retrieval during a later iteration. Based on this auxiliary information, the computer system can make decisions that can greatly reduce the amount of computation required by the training process. The auxiliary information allows the computer system to diagnose and fix problems that occur during the training process on a module-by-module and/or datum-by-datum basis.
Computer systems and methods generate data examples by training, through machine learning, a data generator with a training objective to produce a data example for a specific value of R, where R is value related to S1(x) and S2(x), where, for a data example, x, generated by the data generator, S1(x) is a likelihood that the data example x is in a first class of a first selected data example and S2(x) is a likelihood that the data example x is in a second class of a second selected data example. S1(x) and S2(x) are determined by a discriminator that is trained through machine learning to discriminate between the first and second classes. After training the data generator, the data generator generates a synthetic data example for each of multiple specific values of R.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
Computer systems and computer-implemented methods modify a machine learning network, such as a deep neural network, to introduce judgment to the network. A “combining” node is added to the network, to thereby generate a modified network, where activation of the combining node is based, at least in part, on output from a subject node of the network. The computer system then trains the modified network by, for each training data item in a set of training data, performing forward and back propagation computations through the modified network, where the backward propagation computation through the modified network comprises computing estimated partial derivatives of an error function of an objective for the network, except that the combining node selectively blocks back-propagation of estimated partial derivatives to the subject node, even though activation of the combining node is based on the activation of the subject node.
Computer systems and methods cooperatively train multiple generators and a classifier. Cooperative training includes: training, through machine learning, the multiple generators such that each generator is trained according to a first objective to output examples of a designated classification category; training, through machine learning, the classifier to determine, for each generated by the multiple generators, which of the multiple generators generated the example; and back-propagating partial derivatives of an error cost function from the classifier to the multiple generators.
G06N 7/00 - Agencements informatiques fondés sur des modèles mathématiques spécifiques
G06K 9/62 - Méthodes ou dispositions pour la reconnaissance utilisant des moyens électroniques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06F 17/18 - Opérations mathématiques complexes pour l'évaluation de données statistiques
G06F 12/0815 - Protocoles de cohérence de mémoire cache
42.
Asynchronous agents with learning coaches and structurally modifying deep neural networks without performance degradation
Methods and computer systems improve a trained base deep neural network by structurally changing the base deep neural network to create an updated deep neural network, such that the updated deep neural network has no degradation in performance relative to the base deep neural network on the training data. The updated deep neural network is subsequently training. Also, an asynchronous agent for use in a machine learning system comprises a second machine learning system ML2 that is to be trained to perform some machine learning task. The asynchronous agent further comprises a learning coach LC and an optional data selector machine learning system DS. The purpose of the data selection machine learning system DS is to make the second stage machine learning system ML2 more efficient in its learning (by selecting a set of training data that is smaller but sufficient) and/or more effective (by selecting a set of training data that is focused on an important task). The learning coach LC is a machine learning system that assists the learning of the DS and ML2. Multiple asynchronous agents could also be in communication with each others, each trained and grown asynchronously under the guidance of their respective learning coaches to perform different tasks.
Methods and computer systems improve a trained base deep neural network by structurally changing the base deep neural network to create an updated deep neural network, such that the updated deep neural network has no degradation in performance relative to the base deep neural network on the training data. The updated deep neural network is subsequently training. Also, an asynchronous agent for use in a machine learning system comprises a second machine learning system ML2 that is to be trained to perform some machine learning task. The asynchronous agent further comprises a learning coach LC and an optional data selector machine learning system DS. The purpose of the data selection machine learning system DS is to make the second stage machine learning system ML2 more efficient in its learning (by selecting a set of training data that is smaller but sufficient) and/or more effective (by selecting a set of training data that is focused on an important task). The learning coach LC is a machine learning system that assists the learning of the DS and ML2. Multiple asynchronous agents could also be in communication with each others, each trained and grown asynchronously under the guidance of their respective learning coaches to perform different tasks.
A machine learning system includes a coach machine learning system that uses machine learning to help a student machine learning system learn its system. By monitoring the student learning system, the coach machine learning system can learn (through machine learning techniques) “hyperparameters” for the student learning system that control the machine learning process for the student learning system. The machine learning coach could also determine structural modifications for the student learning system architecture. The learning coach can also control data flow to the student learning system.
A diverse set of neural networks are trained to be individually robust against adversarial attacks and diverse in a manner that decreases the ability of an adversarial example to fool the full diverse set. The systems/methods use a diversity criterion that is specialized for measuring diversity in response to adversarial attacks rather than diversity in the classification results. Also, one or more networks can be trained that are less robust to adversarial attacks to use as a diagnostic to detect the presence of an adversarial attack. Also, node-to-node relation regularization links can be used to train diverse networks that are randomly selected from a family of diverse networks with exponentially many members.
G06N 20/20 - Techniques d’ensemble en apprentissage automatique
G06F 21/57 - Certification ou préservation de plates-formes informatiques fiables, p. ex. démarrages ou arrêts sécurisés, suivis de version, contrôles de logiciel système, mises à jour sécurisées ou évaluation de vulnérabilité
A machine learning system includes a coach machine learning system that uses machine learning to help a student machine learning system learn its system. By monitoring the student learning system, the coach machine learning system can learn (through machine learning techniques) “hyperparameters” for the student learning system that control the machine learning process for the student learning system. The machine learning coach could also determine structural modifications for the student learning system architecture. The learning coach can also control data flow to the student learning system.
Computer systems and computer-implemented methods train a neural network, by:
(a) computing for each datum in a set of training data, activation values for nodes in the neural network and estimates of partial derivatives of an objective function for the neural network for the nodes in the neural network; (b) selecting a target node of the neural network and/or a target datum in the set of training data; (c) selecting a target-specific improvement model for the neural network, wherein the target-specific improvement model, when added to the neural network, improves performance of the neural network for the target node and/or the target datum, as the case may be; (d) training the target-specific improvement model; (e) merging the target-specific improvement model with the neural network to form an expanded neural network; and (f) training the expanded neural network.
Systems and methods analyze training of a first machine learning system with a second machine learning system. The first machine learning system comprises a neural network with a first inner layer node. The method includes connecting the first machine learning system to an input of the second machine learning system. The second machine learning system comprises a second objective function for analyzing an internal characteristic of the first machine learning system and which is different from a first objective function for the first machine learning system. The method further includes providing a training data item to the first machine learning system, collecting internal characteristic data from the first inner layer node of the first machine learning system associated with the internal characteristic, computing partial derivatives of the first objective function through the first machine learning system with respect to the training data item, and computing partial derivatives of the second objective function through both the second machine learning system and the first machine learning system with respect to the collected internal characteristic data.
Data-dependent node-to-node knowledge sharing to increase the interpretability of the activation pattern of one or more nodes in a neural network, is implemented by a set of knowledge sharing links. Each link may comprise a knowledge providing node or other source P and a knowledge receiving node R. A knowledge sharing link can impose a nodespecific regularization on the knowledge receiving node R to help guide the knowledge receiving node R to have an activation pattern that is more easily interpreted. The specification and training of the knowledge sharing links may be controlled by a cooperative human-AI learning supervisor system in which a human and an artificial intelligence system work cooperatively to improve the interpretability and performance of the client system.
n ensemble members such that each of the ensemble members trains with updates in a different direction from each of the other ensemble members. The ensemble members may also be trained with joint optimization.
A computer-implemented method of training an ensemble machine learning system comprising a plurality of ensemble members. The method includes selecting a shared objective and an objective for each of the ensemble members. The method further includes training each of the ensemble members according to each objective on a training data set, connecting an output of each of the ensemble members to a joint optimization machine learning system to form a consolidated machine learning system, and training the consolidated machine learning system according to the shared objective and the objective for each of the ensemble members on the training data set. The ensemble members can be the same or different types of machine learning systems. Further, the joint optimization machine learning system can be the same or a different type of machine learning system than the ensemble members.
G06N 20/20 - Techniques d’ensemble en apprentissage automatique
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
Machine-learning computer system breaks a neural network into a plurality of modules and tracks the training process module-by-module and datum-by-datum, recording auxiliary information during one iteration of the training process for retrieval during a later iteration. Based on this auxiliary information, the computer system can make decisions that can greatly reduce the amount of computation required by the training process. The auxiliary information allows the computer system to diagnose and fix problems that occur during the training process on a module-by-module and/or datum-by-datum basis.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06N 3/06 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
Computer-based systems and methods guide the learning of features in middle layers of a deep neural network. The guidance can be provided by aligning sets of nodes or entire layers in a network being trained with sets of nodes in a reference system. This guidance facilitates the trained network to more efficiently learn features learned by the reference system using fewer parameters and with faster training. The guidance also enables training of a new system with a deeper network, i.e., more layers, which tend to perform better than shallow networks. Also, with fewer parameters, the new network has fewer tendencies to overfit the training data.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
Computer systems and computer-implemented methods modify a machine learning network, such as a deep neural network, to introduce judgment to the network. A "combining" node is added to the network, to thereby generate a modified network, where activation of the combining node is based, at least in part, on output from a subject node of the network. The computer system then trains the modified network by, for each training data item in a set of training data, performing forward and back propagation computations through the modified network, where the backward propagation computation through the modified network comprises computing estimated partial derivatives of an error function of an objective for the network, except that the combining node selectively blocks back-propagation of estimated partial derivatives to the subject node, even though activation of the combining node is based on the activation of the subject node.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
A computer system uses a pool of predefined functions and pre-trained networks to accelerate the process of building a large neural network or building a combination of (i) an ensemble of other machine learning systems with (ii) a deep neural network. Copies of a predefined function node or network may be placed in multiple locations in a network being built. In building a neural network using a pool of predefined networks, the computer system only needs to decide the relative location of each copy of a predefined network or function. The location may be determined by (i) the connections to a predefined network from source nodes and (ii) the connections from a predefined network to nodes in an upper network. The computer system may perform an iterative process of selecting trial locations for connecting arcs and evaluating the connections to choose the best ones.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
57.
Self-organizing partially ordered networks and soft-tying learned parameters, such as connection weights
Computer-implemented systems and methods soft-tie learned parameters of a neural network(s). The soft-tying comprises: applying a common label to the first and second learned parameters; and as part of the training, and in response to the first and second learned parameters having the common label, applying a regularization penalty to a loss function for the first learned parameter upon a determination that the first learned parameter is different than the second learned parameter. The learned parameters can be connection weights, node biases, and/or parametric model statistics. The application of the regularization penalty can be influenced by a soft-tying hyperparameter.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06K 9/62 - Méthodes ou dispositions pour la reconnaissance utilisant des moyens électroniques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06N 20/20 - Techniques d’ensemble en apprentissage automatique
58.
Self-supervised back propagation for deep learning
A computer-implemented method for analyzing a first neural network via a second neural network according to a differentiable function. The method includes adding a derivative node to the first neural network that receives derivatives associated with a node of the first neural network. The derivative node is connected to the second neural network such that the second neural network can receive the derivatives from the derivative node. The method further includes feeding forward activations in the first neural network for a data item, back propagating a selected differentiable function, providing the derivatives from the derivative node to the second neural network as data, feeding forward the derivatives from the derivative node through the second neural network, and then back propagating a secondary objective through both neural networks. In various aspects, the learned parameters of one or both of the neural networks can be updated according to the back propagation calculations.
Systems and methods analyze and correct the vulnerability of individual nodes in a neural network to changes in the input data. The analysis comprises first changing the activation function of one or more nodes to make them more vulnerable. The vulnerability is then measured based on a norm on the vector of partial derivatives of the network objective evaluated on each training data item. The system is made less vulnerable by splitting the data based on the sign of the partial derivative of the network objective with respect to a vulnerable and training new ensemble members on selected subsets from the data split.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
G06N 20/20 - Techniques d’ensemble en apprentissage automatique
G06K 9/62 - Méthodes ou dispositions pour la reconnaissance utilisant des moyens électroniques
G06N 5/04 - Modèles d’inférence ou de raisonnement
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
H04L 67/142 - Gestion des états de session pour les protocoles sans étatÉtats des sessions de signalisationSignalisation des états de sessionMécanismes de conservation d’état
Computer-implemented, machine-learning systems and methods relate to a neural network having at least two subnetworks, i.e., a first subnetwork and a second subnetwork. The systems and methods estimate the partial derivative(s) of an objective with respect to (i) an output activation of a node in first subnetwork, (ii) the input to the node, and/or (iii) the connection weights to the node. The estimated partial derivative(s) are stored in a data store and provided as input to the second subnetwork. Because the estimated partial derivative(s) are persisted in a data store, the second subnetwork has access to them even after the second subnetwork has gone through subsequent training iterations. Using this information, subnetwork 160 can compute classifications and regression functions that can help, for example, in the training of the first subnetwork.
A deep neural network architecture comprises a stack of strata in which each stratum has its individual input and an individual objective, in addition to being activated from the system input through lower strata in the stack and receiving back propagation training from the system objective back propagated through higher strata in the stack of strata. The individual objective for a stratum may comprise an individualized target objective designed to achieve diversity among the strata. Each stratum may have a stratum support subnetwork with various specialized subnetworks. These specialized subnetworks may comprise a linear subnetwork to facilitate communication across strata and various specialized subnetworks that help encode features in a more compact way, not only to facilitate communication across strata but also to increase interpretability for human users and to facilitate communication with other machine learning systems.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
H04L 67/142 - Gestion des états de session pour les protocoles sans étatÉtats des sessions de signalisationSignalisation des états de sessionMécanismes de conservation d’état
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06F 18/214 - Génération de motifs d'entraînementProcédés de Bootstrapping, p. ex. ”bagging” ou ”boosting”
G06F 18/21 - Conception ou mise en place de systèmes ou de techniquesExtraction de caractéristiques dans l'espace des caractéristiquesSéparation aveugle de sources
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
Computer-based systems and methods add extra terms to the objective function of machine learning systems (e.g., neural networks) in an ensemble for selected items of training data. This selective training is designed to penalize and decrease any tendency for two or more members of the ensemble to make the same mistake on any item of training data, which should result in improved performance of the ensemble in operation.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
Various systems and methods are described herein for improving the aggressive development of machine learning systems. In machine learning, there is always a trade-off between allowing a machine learning system to learn as much as it can from training data and overfitting on the training data. This trade-off is important because overfitting usually causes performance on new data to be worse. However, various systems and methods can be utilized to separate the process of detailed learning and knowledge acquisition and the process of imposing restrictions and smoothing estimates, thereby allowing machine learning systems to aggressively learn from training data, while mitigating the effects of overfitting on the training data.
G06N 7/00 - Agencements informatiques fondés sur des modèles mathématiques spécifiques
G06K 9/62 - Méthodes ou dispositions pour la reconnaissance utilisant des moyens électroniques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06F 17/18 - Opérations mathématiques complexes pour l'évaluation de données statistiques
G06F 12/0815 - Protocoles de cohérence de mémoire cache
A machine learning (ML) system includes a student ML system, a learning coach ML system, and a reference system that generates training data for the student ML system. The learning coach ML system learns to make an enhancement to the student ML system or to its learning process, such as updated hyperparameter or a network structural change, based on training of the student ML system with the training data generated by the reference system. The system may also comprise a learning experimentation system that communicates with the reference system to conduct experiments on the learning of the student learning system. Also, the learning experimentation system can determine a cost function for the learning coach ML system.
A machine learning system includes a coach machine learning system that uses machine learning to help a student machine learning system learn its system. By monitoring the student learning system, the coach machine learning system can learn (through machine learning techniques) “hyperparameters” for the student learning system that control the machine learning process for the student learning system. The machine learning coach could also determine structural modifications for the student learning system architecture. The learning coach can also control data flow to the student learning system.
Computer systems and computer-implemented methods train and/or operate, once trained, a machine-learning system that comprises a plurality of generator-detector pairs. The machine-learning computer system comprises a set of processor cores and computer memory that stores software. When executed by the set of processor cores, the software causes the set of processor cores to implement a plurality of generator-detector pairs, in which: (i) each generator-detector pair comprises a machine-learning data generator and a machine-learning data detector; and (ii) each generator-detector pair is for a corresponding cluster of data examples respectively, such that, for each generator-detector pair, the generator is for generating data examples in the corresponding cluster and the detector is for detecting whether data examples are within the corresponding cluster.
G06N 7/00 - Agencements informatiques fondés sur des modèles mathématiques spécifiques
G06K 9/62 - Méthodes ou dispositions pour la reconnaissance utilisant des moyens électroniques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06F 17/18 - Opérations mathématiques complexes pour l'évaluation de données statistiques
G06F 12/0815 - Protocoles de cohérence de mémoire cache
74.
Estimating the amount of degradation with a regression objective in deep learning
Computer systems and computer-implemented methods train a machine-learning regression system. The method comprises the step of generating, with a machine-learning generator, output patterns; distorting the output patterns of the generator by a scale factor to generate distorted output patterns; and training the machine-learning regression system to predict the scaling factor, where the regression system receives the distorted output patterns as input and learns and the scaling factor is a target value for the regression system. The method may further comprise, after training the machine-learning regression system, training a second machine-learning generator by back propagating partial derivatives of an error cost function from the regression system to the second machine-learning generator and training the second machine-learning generator using stochastic gradient descent.
G06F 15/18 - dans lesquels un programme est modifié en fonction de l'expérience acquise par le calculateur lui-même au cours d'un cycle complet; Machines capables de s'instruire (systèmes de commande adaptatifs G05B 13/00;intelligence artificielle G06N)
G06N 7/00 - Agencements informatiques fondés sur des modèles mathématiques spécifiques
G06K 9/62 - Méthodes ou dispositions pour la reconnaissance utilisant des moyens électroniques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06F 17/18 - Opérations mathématiques complexes pour l'évaluation de données statistiques
G06F 12/0815 - Protocoles de cohérence de mémoire cache
Computer systems and methods generate a stochastic categorical autoencoder learning network (SCAN). The SCAN is trained to have an encoder network that outputs, subject to one or more constraints, parameters for parametric probability distributions of sample random variables from input data. The parameters comprise measures of central tendency and measures of dispersion. The one or more constraints comprise a first constraint that constrains a measure of a magnitude of a vector of the measures of central tendency as compared to a measure of a magnitude of a vector of the measures of dispersion. Thereafter, the sample random variables are generated from the parameters and a decoder is trained to output the input data from the sample random variables.
G06N 7/00 - Agencements informatiques fondés sur des modèles mathématiques spécifiques
G06K 9/62 - Méthodes ou dispositions pour la reconnaissance utilisant des moyens électroniques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06F 17/18 - Opérations mathématiques complexes pour l'évaluation de données statistiques
G06F 12/0815 - Protocoles de cohérence de mémoire cache
Computer-implemented, machine-learning systems and methods relate to a combination of neural networks. The systems and methods train the respective member networks both (i) to be diverse and yet (ii) according to a common, overall objective. Each member network is trained or retrained jointly with all the other member networks, including member networks that may not have been present in the ensemble when a member is first trained.
Machine-learning data generators use an additional objective to avoid generating data that is too similar to any previously known data example. This prevents plagiarism or simple copying of existing data examples, enhancing the ability of a generator to usefully generate novel data. A formulation of generative adversarial network (GAN) learning as the mixed strategy minimax solution of a zero-sum game solves the convergence and stability problem of GANs learning, without suffering mode collapse.
G06N 7/00 - Agencements informatiques fondés sur des modèles mathématiques spécifiques
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06K 9/62 - Méthodes ou dispositions pour la reconnaissance utilisant des moyens électroniques
G06F 17/18 - Opérations mathématiques complexes pour l'évaluation de données statistiques
G06F 12/0815 - Protocoles de cohérence de mémoire cache
A machine learning system includes a coach machine learning system that uses machine learning to help a student machine learning system learn its system. By monitoring the student learning system, the coach machine learning system can learn (through machine learning techniques) “hyperparameters” for the student learning system that control the machine learning process for the student learning system. The machine learning coach could also determine structural modifications for the student learning system architecture. The learning coach can also control data flow to the student learning system.
Systems and methods improve the performance of a network that has converged such that the gradient of the network and all the partial derivatives are zero (or close to zero) by splitting the training data such that, on each subset of the split training data, some nodes or arcs (i.e., connections between a node and previous or subsequent layers of the network) have individual partial derivative values that are different from zero on the split subsets of the data, although their partial derivatives averaged over the whole set of training data is close to zero. The present system and method can create a new network by splitting the candidate nodes or arcs that diverge from zero and then trains the resulting network with each selected node trained on the corresponding cluster of the data. Because the direction of the gradient is different for each of the nodes or arcs that are split, the nodes and their arcs in the new network will train to be different. Therefore, the new network is not at a stationary point.
Methods and computer systems improve a trained base deep neural network by structurally changing the base deep neural network to create an updated deep neural network, such that the updated deep neural network has no degradation in performance relative to the base deep neural network on the training data. The updated deep neural network is subsequently training. Also, an asynchronous agent for use in a machine learning system comprises a second machine learning system ML2 that is to be trained to perform some machine learning task. The asynchronous agent further comprises a learning coach LC and an optional data selector machine learning system DS. The purpose of the data selection machine learning system DS is to make the second stage machine learning system ML2 more efficient in its learning (by selecting a set of training data that is smaller but sufficient) and/or more effective (by selecting a set of training data that is focused on an important task). The learning coach LC is a machine learning system that assists the learning of the DS and ML2. Multiple asynchronous agents could also be in communication with each others, each trained and grown asynchronously under the guidance of their respective learning coaches to perform different tasks.
A computer-implemented method for analyzing a first neural network via a second neural network according to a differentiable function. The method includes adding a derivative node to the first neural network that receives derivatives associated with a node of the first neural network. The derivative node is connected to the second neural network such that the second neural network can receive the derivatives from the derivative node. The method further includes feeding forward activations in the first neural network for a data item, back propagating a selected differentiable function, providing the derivatives from the derivative node to the second neural network as data, feeding forward the derivatives from the derivative node through the second neural network, and then back propagating a secondary objective through both neural networks. In various aspects, the learned parameters of one or both of the neural networks can be updated according to the back propagation calculations.
A deep neural network architecture comprises a stack of strata in which each stratum has its individual input and an individual objective, in addition to being activated from the system input through lower strata in the stack and receiving back propagation training from the system objective back propagated through higher strata in the stack of strata. The individual objective for a stratum may comprise an individualized target objective designed to achieve diversity among the strata. Each stratum may have a stratum support subnetwork with various specialized subnetworks. These specialized subnetworks may comprise a linear subnetwork to facilitate communication across strata and various specialized subnetworks that help encode features in a more compact way, not only to facilitate communication across strata but also to increase interpretability for human users and to facilitate communication with other machine learning systems.
A computer system uses a pool of predefined functions and pre-trained networks to accelerate the process of building a large neural network or building a combination of (i) an ensemble of other machine learning systems with (ii) a deep neural network. Copies of a predefined function node or network may be placed in multiple locations in a network being built. In building a neural network using a pool of predefined networks, the computer system only needs to decide the relative location of each copy of a predefined network or function. The location may be determined by (i) the connections to a predefined network from source nodes and (ii) the connections from a predefined network to nodes in an upper network. The computer system may perform an iterative process of selecting trial locations for connecting arcs and evaluating the connections to choose the best ones.
G06N 3/04 - Architecture, p. ex. topologie d'interconnexion
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
A computer-implemented method of training an ensemble machine learning system comprising a plurality of ensemble members. The method includes selecting a shared objective and an objective for each of the ensemble members. The method further includes training each of the ensemble members according to each objective on a training data set, connecting an output of each of the ensemble members to a joint optimization machine learning system to form a consolidated machine learning system, and training the consolidated machine learning system according to the shared objective and the objective for each of the ensemble members on the training data set. The ensemble members can be the same or different types of machine learning systems. Further, the joint optimization machine learning system can be the same or a different type of machine learning system than the ensemble members.
A multi-stage machine learning and recognition system comprises multiple individual machine learning systems arranged in multiple stages, where data is passed from a machine learning system in one stage to one or more machine learning systems in a subsequent, higher-level stage of the structure according to the logic of the machine learning system. The multi-stage machine learning system can be arranged in a final stage and one or more non-final stages, where the one or more non-final stages direct data generally towards a selected one or more machine learning systems within the final stage, but less than all of the machine learning systems in the final stage. The multi-stage machine learning system can additionally include a learning coach and data management system, which is configured to control the distribution of data throughout the multi-stage structure of machine learning systems by observing the internal state of the structure.
G10L 15/16 - Classement ou recherche de la parole utilisant des réseaux neuronaux artificiels
G10L 15/02 - Extraction de caractéristiques pour la reconnaissance de la paroleSélection d'unités de reconnaissance
G10L 15/22 - Procédures utilisées pendant le processus de reconnaissance de la parole, p. ex. dialogue homme-machine
G10L 25/18 - Techniques d'analyse de la parole ou de la voix qui ne se limitent pas à un seul des groupes caractérisées par le type de paramètres extraits les paramètres extraits étant l’information spectrale de chaque sous-bande
Systems and methods for analyzing a first machine learning system via a second machine learning system. The first machine learning system comprising a first objective function. The method includes connecting the first machine learning system to an input of the second machine learning system, which includes a second objective function for analyzing an internal characteristic of the first machine learning system. The method further includes providing a data item to the first machine learning system, collecting internal characteristic data from the first machine learning system associated with the internal characteristic, computing partial derivatives of the first objective function through the first machine learning system with respect to the data item, and computing partial derivatives of the second objective function through both the second machine learning system and the first machine learning system with respect to the collected internal characteristic data.
Computer-implemented systems and methods build and train an ensemble of machine learning systems to be robust against adversarial attacks by employing a probabilistic mixed strategy with the property that, even if the adversary knows the architecture and parameters of the machine learning system, any adversarial attack has an arbitrarily low probability of success.
Computer-implemented systems and methods build ensembles for deep learning through parallel data splitting by creating and training an ensemble of up to 2n ensemble members based on a single base network and a selection of n network elements. The ensemble members are created by the "blasting" process, in which training data are selected for each of the up to 2n ensemble members such that each of the ensemble members trains with updates in a different direction from each of the other ensemble members. The ensemble members may also be trained with joint optimization.
Systems and methods analyze and correct the vulnerability of individual nodes in a neural network to changes in the input data. The analysis comprises first changing the activation function of one or more nodes to make them more vulnerable. The vulnerability is then measured based on a norm on the vector of partial derivatives of the network objective evaluated on each training data item. The system is made less vulnerable by splitting the data based on the sign of the partial derivative of the network objective with respect to a vulnerable and training new ensemble members on selected subsets from the data split.
G06F 11/00 - Détection d'erreursCorrection d'erreursContrôle de fonctionnement
G06F 11/34 - Enregistrement ou évaluation statistique de l'activité du calculateur, p. ex. des interruptions ou des opérations d'entrée–sortie
G06F 15/173 - Communication entre processeurs utilisant un réseau d'interconnexion, p. ex. matriciel, de réarrangement, pyramidal, en étoile ou ramifié
G06F 17/00 - Équipement ou méthodes de traitement de données ou de calcul numérique, spécialement adaptés à des fonctions spécifiques
G06F 17/18 - Opérations mathématiques complexes pour l'évaluation de données statistiques
90.
FORWARD PROPAGATION OF SECONDARY OBJECTIVE FOR DEEP LEARNING
Computer systems and methods optimize a secondary objective function in the training of a multi-layer feed-forward neural network in which the secondary objective is a function of the partial derivatives of the primary objective function. Optimizing this secondary objective function comprises computing derivatives of functions of the partial derivatives computed during the back-propagation computation in a third stage of computation before the parameter update. This third stage of computation proceeds in the reverse direction from the direction of the back propagation computation. That is, the third stage of computation proceeds forwards through the network, computing derivatives of the secondary objective function based on the chain rule of calculus. The secondary objective may be used to make the neural network more robust against deviations in the input values from their normal values.
Computer-implemented, machine-learning systems and methods relate to a neural network having at least two subnetworks, i.e., a first subnetwork and a second subnetwork. The systems and methods estimate the partial derivative(s) of an objective with respect to (i) an output activation of a node in first subnetwork, (ii) the input to the node, and/or (iii) the connection weights to the node. The estimated partial derivative(s) are stored in a data store and provided as input to the second subnetwork. Because the estimated partial derivative(s) are persisted in a data store, the second subnetwork has access to them even after the second subnetwork has gone through subsequent training iterations. Using this information, subnetwork 160 can compute classifications and regression functions that can help, for example, in the training of the first subnetwork.
G06F 15/18 - dans lesquels un programme est modifié en fonction de l'expérience acquise par le calculateur lui-même au cours d'un cycle complet; Machines capables de s'instruire (systèmes de commande adaptatifs G05B 13/00;intelligence artificielle G06N)
G06F 19/24 - pour l'apprentissage automatique, l'exploration de données ou les bio statistiques, p.ex. détection de motifs, extraction de connaissances, extraction de règles, corrélation, agrégation ou classification
A system and method for controlling a nodal network. The method includes estimating an effect on the objective caused by the existence or non-existence of a direct connection between a pair of nodes and changing a structure of the nodal network based at least in part on the estimate of the effect. A nodal network includes a strict partially ordered set, a weighted directed acyclic graph, an artificial neural network, and/or a layered feed-forward neural network.
G06F 15/18 - dans lesquels un programme est modifié en fonction de l'expérience acquise par le calculateur lui-même au cours d'un cycle complet; Machines capables de s'instruire (systèmes de commande adaptatifs G05B 13/00;intelligence artificielle G06N)
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
Computer systems and computer-implemented methods train and/or operate, once trained, a machine-learning system that comprises a plurality of generator-detector pairs. The machine- learning computer system comprises a set of processor cores and computer memory that stores software. When executed by the set of processor cores, the software causes the set of processor cores to implement a plurality of generator-detector pairs, in which: (i) each generator-detector pair comprises a machine-learning data generator and a machine-learning data detector; and (ii) each generator-detector pair is for a corresponding cluster of data examples respectively, such that, for each generator-detector pair, the generator is for generating data examples in the corresponding cluster and the detector is for detecting whether data examples are within the corresponding cluster.
Computer systems and computer-implemented methods train a machine-learning regression system. The method comprises the step of generating, with a machine-learning generator, output patterns; distorting the output patterns of the generator by a scale factor to generate distorted output patterns; and training the machine-learning regression system to predict the scaling factor, where the regression system receives the distorted output patterns as input and learns and the scaling factor is a target value for the regression system. The method may further comprise, after training the machine-learning regression system, training a second machine-learning generator by back propagating partial derivatives of an error cost function from the regression system to the second machine-learning generator and training the second machine-learning generator using stochastic gradient descent.
G06F 15/18 - dans lesquels un programme est modifié en fonction de l'expérience acquise par le calculateur lui-même au cours d'un cycle complet; Machines capables de s'instruire (systèmes de commande adaptatifs G05B 13/00;intelligence artificielle G06N)
95.
ROBUST AUTO-ASSOCIATIVE MEMORY WITH RECURRENT NEURAL NETWORK
Computer systems and computer-implemented methods recursively train a content- addressable auto-associative memory such that: (i) the content addressable auto-associative memory system is trained to produce an output pattern for each of the input examples; and (ii) a quantity of the learned parameters for the content-addressable auto-associative memory is equal to the number of input variables times a quantity that is independent of the number of input variables. The quantity of learned parameters for the content-addressable auto- associative memory system can be varied based on the number of input examples to be learned.
G06F 15/18 - dans lesquels un programme est modifié en fonction de l'expérience acquise par le calculateur lui-même au cours d'un cycle complet; Machines capables de s'instruire (systèmes de commande adaptatifs G05B 13/00;intelligence artificielle G06N)
Machine-learning data generators use an additional objective to avoid generating data that is too similar to any previously known data example. This prevents plagiarism or simple copying of existing data examples, enhancing the ability of a generator to usefully generate novel data. A formulation of generative adversarial network (GAN) learning as the mixed strategy minimax solution of a zero-sum game solves the convergence and stability problem of GANs learning, without suffering mode collapse.
Various systems and methods are described herein for improving the aggressive development of machine learning systems. In machine learning, there is always a trade-off between allowing a machine learning system to learn as much as it can from training data and overfitting on the training data. This trade-off is important because overfitting usually causes performance on new data to be worse. However, various systems and methods can be utilized to separate the process of detailed learning and knowledge acquisition and the process of imposing restrictions and smoothing estimates, thereby allowing machine learning systems to aggressively learn from training data, while mitigating the effects of overfitting on the training data.
Computer-implemented, machine-learning systems and methods relate to a combination of neural networks. The systems and methods train the respective member networks both (i) to be diverse and yet (ii) according to a common, overall objective. Each member network is trained or retrained jointly with all the other member networks, including member networks that may not have been present in the ensemble when a member is first trained.
Computer systems and methods generate a stochastic categorical autoencoder learning network (SCAN). The SCAN is trained to have an encoder network that outputs, subject to one or more constraints, parameters for parametric probability distributions of sample random variables from input data. The parameters comprise measures of central tendency and measures of dispersion. The one or more constraints comprise a first constraint that constrains a measure of a magnitude of a vector of the measures of central tendency as compared to a measure of a magnitude of a vector of the measures of dispersion. Thereafter, the sample random variables are generated from the parameters and a decoder is trained to output the input data from the sample random variables.
Computer-based systems and methods add extra terms to the objective function of machine learning systems (e.g., neural networks) in an ensemble for selected items of training data. This selective training is designed to penalize and decrease any tendency for two or more members of the ensemble to make the same mistake on any item of training data, which should result in improved performance of the ensemble in operation.
G06F 19/24 - pour l'apprentissage automatique, l'exploration de données ou les bio statistiques, p.ex. détection de motifs, extraction de connaissances, extraction de règles, corrélation, agrégation ou classification