In some aspects, a computing system can generate and optimize a hybrid machine learning model for risk assessment based on predictor variables associated with a target entity. The hybrid machine learning model can be trained using training vectors with sets of training predictor variables and training outputs corresponding to the respective sets of training predictor variables. The predictor variables associated with the target entity may include unknown values and the training predictor variables or trainings output may also include unknown values. Additionally, the computing system can generate explanatory data for the target entity to indicate relationships between changes in the risk indicator and changes in the predictor variables associated with the target entity. The risk indicator and the explanatory data can be used in controlling access of the target entity to interactive computing environments.
A system can identify entities and verify entity attributes by leveraging the ability of trained large language model (LLM) to extract entity information from data instances of one or more data sources. The data instances may be, for example, writings such as news articles. The writings may include information about entities and/or events in which the entities were involved. The system may generate a pair of input prompts for the LLM to cause the LLM to process a data instance(s). Each prompt may include a deliberately false data item and the LLM may be caused to generate a response for each prompt. The validity of the information in the responses can be verified by determining that false data items in both responses match. The system can configure another computing system to utilize information about a given entity identified in the data instance and included in the responses.
In some aspects, a computing system can receive a set of candidate records and a query record, and for each record, tokenize the record into a set of tokens and transform each token of the set of tokens into a respective bloom filter, including a query bloom filter, and a set of candidate bloom filters. The system can generate a query bloom filter – candidate bloom filter pair for each candidate bloom filter in the set of candidate bloom filters, and for each query bloom filter – candidate bloom filter pair, identify a distance between the query bloom filter and the candidate bloom filter. The computing system can then identify a selected set of candidate records based on the selected set of candidate records falling under a distance threshold among the query bloom filter – candidate bloom filter pairs, and link the query record to the selected set of candidate records.
A system can generate content recommendations and facilitate interactions using machine-learning. The system can receive a request from a provider entity. The system can receive entity data and interaction data associated with a target entity. The system can generate at least a first graph structure and a second graph structure. The system can generate a linked graph structure based on the first graph structure and the second graph structure. The system can determine among a plurality of operations, one or more target operations to perform on data included in the linked graph structure. The system can execute using a trained machine-learning model, the target operations to generate a content recommendation for facilitating an interaction. The system can provide a responsive message based on the content recommendation usable to facilitate the interaction.
A method can include receiving a request associated with an interaction involving a target entity. The method can include receiving identity data about the target entity and historical data about historical interactions associated with the target entity. The method can include providing the identity data to a first machine-learning model to generate first output. The method can include applying the historical data to rules to generate rule outcomes. The method can include providing the rule outcomes to a second machine-learning model to generate second output. The method can include generating a recommendation including a response to the request. The method can include providing a responsive message including a command executable to automatically control the response to the request.
According to one example, a non-transitory computer-readable storage medium having program code executable by a processing device to perform operations is described. The operations can include receiving, from a third-party device, a digital interaction inquiry associated with a digital interaction and retrieving interaction data associated with the digital interaction. The digital interaction can include a transfer of electronic resources from a third-party associated with the third-party device to a separate entity. The operations can include retrieving interaction data associated with the digital interaction and applying the interaction data to a machine-learning model trained on dispute training data to generate a dispute score. The dispute score can represent a likelihood of successfully disputing the digital interaction and reversing the transfer of the electronic resources. The operations can include transmitting the dispute score to the third-party device, to cause the third-party to respond to the digital interaction.
Systems and methods for predicting future risk for a target entity are provided. A risk assessment system receives historical risk assessment data of the target entity and identifies a target cluster that matches the historical risk assessment data. The target cluster is identified from a group of clusters determined using high dimensional clustering based on risk assessment data of a set of entities. The risk assessment system identifies a set of nearest neighbors of the target cluster and determines a prediction of future risk for the target entity based on the target cluster and the set of nearest neighbors. The risk assessment system transmits a responsive message, which can include the prediction of future risk, to a remote computing device for use in controlling access of the target entity to one or more interactive computing environments.
A system can generate a risk assessment associated with an identity element used in interactions associated with a user entity. For example, the system can receive historical data related an identity element associated with a user entity, the identity element used in a set of interactions associated with the user entity. The system can generate a binomial distribution of the historical data associated with the identity element. The system can determine, based at least in part on the binomial distribution of the historical data, a risk indicator associated with the identity element. The system can control, based at least in part on the risk indicator associated with the identity element, an interaction involving a target entity and the user entity using the identity element.
A device stacking detection computing system receives single-provider data from a service provider computing system associated with a telecommunications service provider. The single-provider data describes a request for a digital device, such as a consumer request to receive a digital device for a new account. Based on the single-provider data, the device stacking detection computing system determines multi-provider data associated with additional telecommunications service providers. Based on a combination of the single-provider data and the multi-provider data, the device stacking detection computing system identifies a risk level for the request, e.g., a risk of device stacking fraud. The device stacking detection computing system provides, to the service provider computing system, data indicating the risk level. In some cases, the device stacking detection computing system provides alert data to additional service provider computing systems associated with the additional telecommunications service providers.
In some aspects, a computing system can improve a machine learning model for risk assessment by removing or reducing bias in the machine learning model. The training process for the machine learning model can include training the machine learning model using training samples, obtaining data for a protected attribute, and calculating a bias metric using the data for the protected attribute and data obtained from the trained machine learning model. Based on the bias metric, bias associated with the machine learning model can be detected. The machine learning model can be modified based on the detected bias and re-trained. The re-trained machine learning model can be used to predict a risk indicator for a target entity. The predicted risk indicator can be transmitted to a remote computing device and be used for controlling access of the target entity to one or more interactive computing environments.
Systems and methods for obfuscating sensitive data records by aggregating the data based on data attributes are provided. Each sensitive data record can contain at least one sensitive attribute. A data protection system can receive a request to access at least a portion of the sensitive data records to which access is restricted. The data protection system can transform the sensitive data records using a data transformation model. The data transformation model can be determined based on a targeted use of aggregated data. The data protection system can group the sensitive data records into aggregation segments by executing a data segmentation model. The data protection system can generate the aggregated data by combining individual sensitive attributes of the sensitive data records based on the aggregation segments.
In some aspects, a computing system can train a machine-learning model to analyze a graph database for risk assessment. The computing system can use the machine-learning model to identify a risk indicator for a target component of one or more interactive computing environments. The graph database can include a set of nodes where each node represents a respective infrastructure service of one or more infrastructure services and a set of edges connecting individual nodes of the set of nodes. The computing system can generate the risk indicator for the target component based on an output of the machine-learning model The computing system additionally can output a graphical user interface including at least the risk indicator for use in controlling access to the one or more infrastructure services.
In one example, a method is described, including receiving candidate data associated with an entity, where the data includes both a first set of entity attributes related to fraud and a second set of attributes associated with entity dispute patterns. The method includes generating, by a first model, a risk score for the entity by evaluating the extent to which a threshold number of attribute-based rules are triggered within the first set of anomalous behavior patterns-related attributes. The operations include generating, by a second model, a separate risk score for the entity by clustering the second set of dispute-related attributes. The method includes generating an aggregate risk score by combining the outputs of the first and second models to then provide an assessment of the entity's risk profile.
G06Q 20/40 - Authorisation, e.g. identification of payer or payee, verification of customer or shop credentialsReview and approval of payers, e.g. check of credit lines or negative lists
14.
CONTROLLING ACCESS USING MULTIPLE DATA SOURCES WITH VARYING AVAILABILITY
In some aspects, a computing system can train a machine learning (ML) model for risk assessment. Once trained, the ML model can determine a risk indicator for a target entity that indicates a level of risk associated with the target entity. Training the ML model can include: using a foundational model pre-trained to predict multiple outcomes to compute the set of common features; and training the machine learning model using the computed set of common features as training inputs.
In some aspects, a graph-merge-split computing system is provided. The graph-merge-split computing system is capable of identifying candidate records for merging and generate an entity level graph from the list of candidate records comprising nodes with matching score edges. A graph-merge process is performed for each connected component which includes determining the matching score for each edge. Edges falling below a threshold are removed. A maximal graph matching algorithm is applied the updated graph where the maximal graph matching algorithm identifies paired nodes. For each of the paired nodes, the graph-merge process then determines a clique score and in response to determining the clique score exceeds a delta-clique constraint, applies a graph-splitter process to the paired nodes prior to forming the merged entity. The graph-merge process is repeated until no merge occurs. Once all merges have been performed, an output graph is generated.
In one example, a computer-implemented method includes receiving, by a processor, a query comprising an entity name, where the entity name is representative of an entity. The method includes retrieving, by the processor and from one or more repositories, an entity data set comprising records associated with the entity and prompt engineering a large language model (LLM) with the entity data set. The method includes generating, by the processor and using the prompt engineered LLM, a vertical prediction based at least in part on the entity name and assigning the vertical prediction to an entity record corresponding to the entity. The method also includes storing, by the processor, the entity record in an enriched entity data set.
In some aspects, a computing system can include an embedding-based search system for retrieving records. For example, the system can store one or more personally identifiable information ("PII") data sets where each PII record corresponds to a respective entity and generate, by a classification model, PII vector embeddings for one or more of the PII records. The system can receive one or more queries and generate, by a classification machine learning mode, PII vector embeddings from a generate, by the classification model, a query vector embedding corresponding to one of the one or more queries. The system may then determine a similarity score by determining a distance between the query vector embedding and the one or more of the PII vector embeddings.
A power graph convolutional network (PGCN) can be used for explainable machine learning. For example, a computing device can determine, using a PGCN, a risk indicator for a target entity from predictor variables associated with the target entity. The PGCN includes a convolutional layer configured to generate modified predictor variables by multiplying each of the predictor variables with a respective interaction factor generated based on an adjacency weight matrix and the predictor variables. The PGCN also includes a dense layer configured to generate the risk indicator. The training process involves adjusting a set of weights in the adjacency weight matrix and a set of weights in a weight vector of the dense layer based on a loss function of the PGCN. The computing device transmits a responsive message including the risk indicator for use in controlling access of the target entity to one or more interactive computing environments.
A system can receive data that includes a first subset and a second subset relating to a first application and a second application in a cloud computing environment. The system can determine differences between the first subset and the second subset by comparing the first subset and the second subset. The system can generate, by using the differences, quality metrics for a transformation from the first application to the second application in the cloud computing environment. The system can generate, by using historical data that includes historical first subsets and second subsets, a trend that represents a progression of the transformation. The system can determine, by using the differences, the quality metrics, and the trend, code in the second application that is likely to cause the differences. The system can generate a command that is executable in the cloud computing environment to automatically update the code in the second application.
An event restraint-delivery computing system receives event data in a high-volume event stream from source computing systems. For each event data object in the high-volume event data stream, the event restraint-delivery computing system identifies at least one event data object subset associated with a respective recipient computing system. The event restraint-delivery computing system withholds the event data object subset from the respective recipient computing system based on restraint status data for the event data object subset. Based on a change in the respective restraint status data, such as a change indicating that the event data object subset is modified to fulfill a trigger criterion, the event restraint-delivery computing system provides the event data object subset to the respective recipient computing system. In addition, the respective recipient computing system performs a computing function based on event data in the event data object subset.
An allocation computing system includes a simulation agent, a manager agent, and a processor pool of multiple processors that are configured to receive workloads. Based on processor usage data about workloads processed by the processors, the simulator agent calculates dynamic capacity of each processor. A machine-learning model in the simulator agent generates respective prediction data that describes an estimated future capacity for each of the processors. Based on the respective prediction data, the manager agent determines a matching relationship between a particular queued workload and the estimated future capacity for a particular processor. The manager agent allocates the particular queued workload to the corresponding processor.
A method to verify an identity document using a machine learning model. The method may comprise receiving an image of the identity document by a computing system, extracting, the image from the identity document, preprocessing, the image prior to providing the image to the machine learning model and providing the preprocessed image to the machine learning model as input. The machine learning model may analyze the input to detect whether the image was synthetically generated. The machine learning model may comprise a convolution layer configured to extract spatial features from the image; a batch normalization layer configured to stabilize a feature map; and a dense layer for classification. An output of the machine learning model may be generated. The output may be an image verification score indicating a likelihood that the image was synthetically generated. Based on the score, an authentication method or a denial message may be generated.
In some aspects, a computing system can train a machine learning (ML) model for risk assessment. Once trained, the ML model can determine a risk indicator for a target entity that indicates a level of risk associated with the target entity. Training the ML model can include: using a foundational model pre-trained to predict multiple outcomes to compute the set of common features; and training the machine learning model using the computed set of common features as training inputs.
A system can efficiently determine whether an identity is manipulated. The system can receive entity data and interaction data associated with a target entity. The system can determine, based on the entity data and the interaction data, one or more risk signals associated with the target entity using one or more artificial intelligence models. The system can generate a linked graph structure based on a first graph structure and a second graph structure each generated using the entity data and the interaction data. The system can apply the one or more risk signals to the linked graph structure to determine a risk indicator associated with the target entity. The system can provide a responsive message based on the risk indicator. The responsive message can be used to control access of the target entity to an interactive computing environment.
G06F 21/57 - Certifying or maintaining trusted computer platforms, e.g. secure boots or power-downs, version controls, system software checks, secure updates or assessing vulnerabilities
25.
RISK ASSESSMENT TECHNIQUES FOR CONTROLLING ACCESS TO COMPUTING SYSTEMS USING DYNAMICALLY DEPLOYED MACHINE LEARNING MODELS
A system can generate a risk assessment associated with a target entity. The system can access a request for a risk indicator associated with a target entity. The system can also access a model in a test environment and run an ATP tool on the model. The ATP tool can execute a set of deployment steps using the model and determine a completion status of each of the set of deployment steps. The system can deploy the model to a production environment when the set of deployment steps is successful. The system can determine the risk indicator using the model. The system can transmit, to a remote computing device, a responsive message comprising at least the risk indicator to control access of the target entity to one or more interactive computing environments.
The present disclosure relates to methods and systems for training and utilizing a machine- learning model with a Deep Kernel Learning with Gaussian processes (DKL-GP) architecture to handle datasets with missing values. The system can receive a dataset with incomplete data, identify missing values, and process the dataset using the DKL-GP architecture. This can involve generating latent variables, utilizing inducing variables to approximate a Gaussian process, and mapping the latent variables to output predictions with associated uncertainty estimates. The system can optimize model parameters through a training process that leverages Pólya-Gamma data augmentation and Gaussian process inducing points for efficient computation. The trained model can subsequently be used to generate predictions for data records with missing data values, while obviating the need to impute potential values for the missing values, and make decisions based on the predictions.
A system can be used to automatically generate infrastructure using a dynamic DAG. The system can receive an input file that can include indications of computing resources for a cloud computing environment. The system can parse the input file to determine a topology of the computing resources based on indications. The system can generate, based on the topology, a dynamic DAG for generating infrastructure for the cloud computing environment based on the computing resources. Generating the dynamic DAG can include (i) determining a first subset of the computing resources and a second subset of the computing resources, and (ii) configuring the first subset of the computing resources to be processed differently from the second subset of the computing resources.
In some aspects, a computing system can train a risk assessment model, using a training process, for determining a risk indicator. The training process can include: accessing a set of features; determining, using a selector network, a set of selected features and an indicator vector; and training the risk assessment model using the set of selected features and the indicator vector. The computing system can determine the risk indicator for a target entity using the trained risk assessment model. The computing system can transmit, to a remote computing device, a responsive message including at least the risk indicator for use in controlling access of the target entity to one or more interactive computing environments.
Systems and methods for secure data management are provided. An application executing on a client computer receives first entity data. The application generates a hashed value based on a subset of the first entity data and transmits the hashed value to a client-facing intermediate layer. A server computer retrieves the hashed value and a client identifier associated with the client computer. The server computer identifies, using the hashed value and the client identifier, a data key. The server computer retrieves, using the data key, second entity data from a database associated with the server computer and transmits the second entity data to the client computer.
G06F 21/62 - Protecting access to data via a platform, e.g. using keys or access control rules
H04L 9/32 - Arrangements for secret or secure communicationsNetwork security protocols including means for verifying the identity or authority of a user of the system
In some aspects, a computing system can train a machine learning (ML) model for risk assessment. Once trained, the ML model can determine a risk indicator for a target entity that indicates a level of risk associated with the target entity. Training the ML model can include: receiving compliance metadata and the machine learning model; for each feature, comparing an associated usage value with a threshold usage value for each of the features associated with the machine learning model; for each feature, including that feature in a compliant set of training features based on the comparison; generating a training dataset comprising the compliant set of training features; and generating an updated machine learning model by training the machine learning model using the training dataset.
THE UNIVERSITY OF NORTH CAROLINA AT CHAPEL HILL (USA)
Inventor
Tian, Longxiu
Zhao, Tian
Miller, Stephen
Abstract
The present disclosure relates to methods and systems for training and utilizing a machine-learning model with a Deep Kernel Learning with Gaussian processes (DKL-GP) architecture to handle datasets with missing values. The system can receive a dataset with incomplete data, identify missing values, and process the dataset using the DKL-GP architecture. This can involve generating latent variables, utilizing inducing variables to approximate a Gaussian process, and mapping the latent variables to output predictions with associated uncertainty estimates. The system can optimize model parameters through a training process that leverages Pólya-Gamma data augmentation and Gaussian process inducing points for efficient computation. The trained model can subsequently be used to generate predictions for data records with missing data values, while obviating the need to impute potential values for the missing values, and make decisions based on the predictions.
Techniques may include generating a graph data structure from transaction data by: accessing transaction data that includes identifiers; processing the transaction data to obtain a set of identifiers from at least one transaction; and creating the graph having a set of nodes and edges, where each node represents the identifier and each edge connects two identifiers from the same transaction. In addition, the techniques may include processing the graph to merge entity labels by: processing nodes with a common entity label to determine a dominant name for the common entity label, where the dominant name is a most frequent name from the common entity label's nodes; and processing the graph to identify entity labels that share a common dominant name; and merging the entity labels that share the common dominant name. The techniques may include performing an operation with the graph.
G06Q 20/40 - Authorisation, e.g. identification of payer or payee, verification of customer or shop credentialsReview and approval of payers, e.g. check of credit lines or negative lists
36 - Financial, insurance and real estate services
42 - Scientific, technological and industrial services, research and design
Goods & Services
Business consulting, management, and planning services in the field of credit information enabled by artificial intelligence, namely, machine learning, mathematical modeling, advanced statistical techniques, and financial analysis of big data, proprietary data, and alternative data Financial information and advisory services; financial risk assessment services; credit risk management; Financial credit scoring services; all the foregoing enabled by artificial intelligence, namely, machine learning, mathematical modeling, advanced statistical techniques, and financial analysis of big data, proprietary data, and alternative data Application service provider (ASP) featuring software for use in credit application processing, evaluating credit worthiness, risk analysis; Fraud alert services, namely, monitoring consumer credit reports and providing an alert as to any changes therein; all of the foregoing enabled by artificial intelligence, namely, machine learning, mathematical modeling, advanced statistical techniques, and financial analysis of big data, proprietary data, and alternative data
34.
TARGET PERMUTATION TESTING TO GENERATE COMPUTATIONALLY EFFICIENT MODELS
A computing device can train a model on original data and determine first characteristics of the trained model. The computing device can permute a target variable of the original data by randomly shuffling the target variable to obtain a set of permutations of the original data. The computing device can train a permuted model for each permutation of the set of permutations to generate a set of permuted models, determine second characteristics for each permuted model of the set of permuted models, and compare the second characteristics for each permuted model and the first characteristics of the trained model to determine relevance of a set of features associated with the first characteristics to an output of the trained model. The computing device can adjust the trained model to include only a subset of features that are relevant to the output of the trained model.
Systems and methods for facilitating and optimizing enterprise personnel communications using a multimodal artificial intelligence (AI) meeting assistive engine. The multimodal AI meeting assistive engine can receive or otherwise obtain information relative to user communications (e.g., meetings) and other user functions and responsibilities from a number of sources. These sources can include but are not limited to sensors, which may take the form of hardware devices such as cameras and microphones that are commonplace in a typical enterprise setting. From the sensors, the multimodal AI meeting assistive engine can receive signals that are representative of one or more physical characteristics of a user, and can determine if a user physical characteristic, such as a facial expression or a user voice, are abnormal. Abnormal physical characteristics may be modified by the multimodal AI meeting assistive engine so that the abnormal physical characteristics appear normal to virtual third parties.
A system can generate a risk assessment associated with a target entity. The system can receive a request for a risk indicator. At a first time and in a first environment, the system can train a first model using a first training data set. At the first time and in a second environment, the system can train a second model using a second training data set. At a second time, the system can select the second model based on a comparison of metrics associated with each model. The system can deploy the second model to the first environment and can determine the risk indicator using the second model. The system can transmit, to a remote computing device, a responsive message including the risk indicator to control access of the target entity to one or more interactive computing environments.
G06Q 10/0635 - Risk analysis of enterprise or organisation activities
G06F 21/57 - Certifying or maintaining trusted computer platforms, e.g. secure boots or power-downs, version controls, system software checks, secure updates or assessing vulnerabilities
37.
RISK ASSESSMENT TECHNIQUES BASED ON DYNAMIC DATA SELECTION
A system can generate a data source selection from a set of data sources. For example, the system can receive a data selection request indicating a set of data sources. For each data source in a set of data sources, the system can: generate a random data sample, a fraud data sample, and a synthetic data sample; generate a set of metrics based on the random data sample, the fraud data sample, and the synthetic data sample, and based on the set of metrics, determine a data source score. The system can select, a data source from the set of data sources based at least in part on the data source score. The system can also transmit, to a remote computing device, a responsive message including at least an indication of the selected data source.
G06Q 20/40 - Authorisation, e.g. identification of payer or payee, verification of customer or shop credentialsReview and approval of payers, e.g. check of credit lines or negative lists
In some aspects, systems and methods for efficiently clustering a large-scale dataset for improving the construction and training of machine-learning models, such as neural network models, are provided. Clustering can include determining a number of clusters to be generated for the dataset. A dataset used for training a neural network model configured can be clustered into a set of clusters. The clustering can include determining the number of clusters, determining special features for the determined number of clusters, and re-clustering the dataset based on the special features. The neural network can be trained based on training samples selected from the set of clusters. In some aspects, the trained neural network model can be utilized to satisfy risk assessment queries to compute output risk indicators for target entities. The output risk indicator can be used to control access to one or more interactive computing environments by the target entities.
G06Q 20/40 - Authorisation, e.g. identification of payer or payee, verification of customer or shop credentialsReview and approval of payers, e.g. check of credit lines or negative lists
39.
CONSOLIDATION OF DATA SOURCES FOR EXPEDITED VALIDATION OF RISK ASSESSMENT DATA
Systems and methods for consolidating data sources for expediting validation processes are described herein. A user interface can be provided to a first entity. First risk assessment data associated with the first entity can be received via the user interface. Second risk assessment data associated with the first entity can be received. The first risk assessment data and the second risk assessment data can be validated. The first risk assessment data and the second risk assessment data can be output for display on the user interface. A responsive message including the validated first risk assessment data and the validated second risk assessment data can be transmitted to a remote computing device for use in controlling access of the first entity to one or more interactive computing environments.
A system can generate a risk assessment associated with a target entity. For example, the system can receive a request for a risk indicator associated with a target entity. The system can retrieve a header record from a database where the record includes a locator and identity data. The system can query an external database associated with the locator to retrieve event data and entity identity data. The system can determine that an entity associated with the entity identity data is the target entity. The system can determine a risk indicator by applying the event data to an algorithm. The system can also transmit, to a remote computing device, a responsive message including at least the risk indicator for use in controlling access of the target entity to one or more interactive computing environments.
Systems and methods for using machine-learning techniques to provide risk assessment based on multiple sources of data are described herein. Data about an entity can be received, and the data can be authenticated. Integrated risk data about the entity can be received. The integrated risk data can include traditional risk assessment data and nontraditional risk assessment data. An integrated risk assessment value can be determined based on the integrated risk data by aligning a first output from a first risk assessment model and a second output by a second risk assessment model. A responsive message including at least the integrated risk assessment value and associated information for the entity can be transmitted to a remote computing device for use in controlling access of the entity to one or more interactive computing environments.
G06F 21/57 - Certifying or maintaining trusted computer platforms, e.g. secure boots or power-downs, version controls, system software checks, secure updates or assessing vulnerabilities
42.
TECHNIQUES FOR CONTROLLING ACCESS TO COMPUTING SYSTEMS BASED ON A RISK SCORE
A system can generate a risk indicator associated with a target entity. For example, the system can receive a request for a risk indicator associated with a target entity. For each data source in a set of data sources, the system can: retrieve identity data associated with the target entity based on the identity of the target entity; and generate a set of element risk scores and a set of affiliation scores associated with each element of the set of elements. The system can determine an aggregate element risk score and an aggregate element affiliation score. The system can determine the risk indicator by combining the aggregated element risk scores of the set of elements based on a first set of element weights. The system can transmit, to a remote computing device, a responsive message including at least the risk indicator.
A system can generate a trust indicator associated with a target entity. For each data source, the system can: retrieve identity data associated with the target entity based on the identity of the target entity; generate a set of element risk scores and a set of affiliation scores associated with each element of the set of elements. The system can determine an aggregate element risk score and an aggregate element affiliation score. The system can determine a risk score by combining the aggregated element risk scores based on a first set of element weights and an affiliation score by combining the aggregated element affiliation scores based on a second set of weights. The system can transmit a responsive message including at least the trust indicator in which the trust indicator is based on the risk score and the affiliation score.
In some aspects, a verification system can receive a verification query from a verifier computing system for requesting verification of characteristics of an entity involved in an online interaction. The verification query can include a unique identifier (“UID”) of the entity. The verification computing system can query a verification repository in the verification computing system based on the UID. Additionally, the verification computing system can query an external-source cache using the UID. In response to determine a match for the UID in the external-source cache, the verification computing system can request external sensitive data records for the entity from an external source corresponding to the external-source cache. Generating consolidated sensitive data records can involve consolidating the external sensitive data records and internal sensitive data records obtained through querying the verification repository. A verification result, generated using the consolidated sensitive data records, can be transmitted to the verifier computing system.
In some aspects, a compliance computing system can generate an evidence repository including evidence data received via an application programming interface (API) from one or more external systems. The system can generate an evidence repository including evidence data received via an application programming interface (API) from one or more systems. The system can receive a compliance request including a compliance query, the compliance query including a requirement identifier associated with a target requirement. The system can also query the evidence repository based on the requirement identifier to retrieve the evidence data associated with the target requirement. Finally, the system can transmit, via a firewall of the compliance computing system, a response message to an external computing system, wherein the response message includes the retrieved evidence data.
Systems and methods for securely sharing stored personal information using data encryption and biometric authentication techniques are disclosed. In some examples, unique items of personal information of an individual may be obtained and placed in data packets. The data packets may be encrypted and may also be encoded with the identity of the individual. A chain of the encrypted data packets may be created and stored in a data repository. A request from an authorized entity for specific personal information of the individual can be received by the system, and may include consent to the request by the individual, and a private key that may be encoded with the identity of the individual. The system can validate the request, and can thereafter decrypt the encrypted data packet containing the requested personal information using the private key and provide the requested personal information to the entity.
H04L 9/32 - Arrangements for secret or secure communicationsNetwork security protocols including means for verifying the identity or authority of a user of the system
H04L 9/00 - Arrangements for secret or secure communicationsNetwork security protocols
47.
EXPONENTIALLY SMOOTHED CATEGORICAL ENCODING TO CONTROL ACCESS TO A NETWORK RESOURCE
In an example of a method described herein, historical events occurring over a network are detected, and at least one of the historical events is associated with an observed value of a categorical variable. A numerical aggregate value representing the observed value is updated by applying an exponential smoothing function to (i) a prior numerical aggregate value representing prior historical events associated with the observed value and (ii) a count of the historical events associated with the observed value. An event occurring over the network is detected and is associated with the observed value. Features are extracted from the event, where the features include an encoded feature based on the numerical aggregate value to represent the observed value. A predictive model is applied to the features to determine a score representing likelihood of an outcome. Based on the score, access to a resource of the network is controlled.
A method described herein involves various operations directed toward network security. The operations include receiving a request for a risk indicator of a target entity. The risk indicator can indicate a level of risk associated with the target entity based on whether the target entity is associated with a mobile device emulator. The operations include generating the risk indicator by applying the entity data to an emulator detection model trained on a training dataset comprising a corpus of attribute data and interaction data. Finally, the operations include providing, to a remote computing device, a responsive message comprising at least the risk indicator to control access to an interactive computing environment.
A system can generate a risk assessment associated with a target entity. For example, the system can receive a request for a risk indicator associated with a target entity. The system can determine that a data source contains a name associated with the target entity based on extracted text from the data source. The system can identify a sentence containing the name. The system can further determine a sentiment score for the sentence. The system can generate a classification associated with an event included in the sentence. The system can determine a confidence score that the name in the extracted text is associated with the target entity based on attributes associated with the name in the extracted text. The system can transmit, to a remote computing device, a message including the risk indicator based on the classification or sentiment score.
A system can generate a risk assessment associated with a target entity. For example, the system can receive a request for a risk indicator associated with a target entity. For each data source in a set of data sources, the system can: retrieve identity data associated with the target entity based on the identity of the target entity; and generate a set of element scores associated with each element of the set of elements. The system can determine an aggregate element score by combining the data source-level element scores for the set of data sources. The system can determine the risk indicator by combining the aggregated element scores of the set of elements based on a set of element weights. The system can also transmit, to a remote computing device, a responsive message including at least the risk indicator.
Systems and methods for creating predictor variables from unstructured data for prediction models are provided. A variable creation application receives unstructured data and processing the unstructured data to generate processed data. Based on the processed data, the variable creation application generates an attribute pool that contains multiple predictor variables generated by applying natural language processing (NLP) procedures on the processed data. The variable creation application further executes a prediction model on at least the predictor variables in the attribute pool to generate a prediction result. Based on the prediction result, the variable creation application evaluates the predictive power of each of the predictor variables and retains predictor variables that are predictive as input predictor variables for the prediction model.
Systems and methods for automated path-based recommendation for risk mitigation are provided. An entity assessment server, responsive to a request for a recommendation for modifying a current risk assessment score of an entity to a target risk assessment score, accesses an input attribute vector for the entity and clusters of entities defined by historical attribute vectors. The entity assessment server assigns the input attribute vector to a particular cluster and determines a requirement on movement from a first point to a second point in a multi-dimensional space based on the statistics computed from the particular cluster. The first point corresponds to the current risk assessment score and the second point corresponds to the target risk assessment score. The entity assessment server computes an attribute-change vector so that a path defined by the attribute-change vector complies with the requirement and generates the recommendation from the attribute-change vector.
Systems and methods for secure resource management are provided. A secure resource management system includes a resource record repository, such as a secure database or a blockchain, for storing resource records for resources. The resource records contain information of resource providers, information of resource users having a right to obtain resources, and resource transaction histories. Responsive to a request to verify an authorized user of a resource, the secure resource management system further queries the resource record repository, retrieves the resource record, determines the resource user currently having a right to obtain the resource as the authorized user of the resource, and transmits the verification result in response to the request. The verification result identifies the authorized user of the resource and can be used to grant access to the resource by the authorized user.
In some aspects, a computing system can train a machine learning (ML) model for risk assessment using risk assessment training data generated at least in part by a prediction model. Once trained, the ML model can determine a risk indicator for a target entity that indicates a level of risk associated with the target entity. Training the ML model can involve receiving a request to predict a value of an unknown segment of a transaction record used to train the ML model. The computing system can execute the prediction model to predict the value of the unknown segment prior to training the ML model using the transaction record that includes the predicted value generated by the prediction model. Additionally, the computing system can generate explanatory data for the target entity indicating relationships between changes in the risk indicator and changes in the transaction records associated with the target entity.
A system can generate a risk assessment associated with an identity element used in interactions associated with a user entity. For example, the system can receive historical data related an identity element associated with a user entity, the identity element used in a set of interactions associated with the user entity. The system can generate a binomial distribution of the historical data associated with the identity element. The system can determine, based at least in part on the binomial distribution of the historical data, a risk indicator associated with the identity element. The system can control, based at least in part on the risk indicator associated with the identity element, an interaction involving a target entity and the user entity using the identity element.
A telecommunications network server system provides a digital identifier to a user device. The digital identifier may include identification data corresponding to a user of the user device. In addition, the telecommunications network server system receives, from one or more third-party systems, requests to authenticate the user for an electronic transaction with the respective third-party system. The telecommunications network server system provides a unique electronic transaction code to each third-party system. Responsive to receiving from the user device one of the unique electronic transaction codes, the telecommunications network server system provides, to the respective third-party system, authentication of the user.
H04L 9/32 - Arrangements for secret or secure communicationsNetwork security protocols including means for verifying the identity or authority of a user of the system
A system can be used to control reversal of an interaction. The system can receive a request to reverse a previously executed interaction. The request can include data relating to a previously executed interaction that may be associated with a. target entity. The system can generate, using an artificial intelligence model that includes a generative artificial intelligence model, risk signals based on the request and the data. The system can determine, based on the risk signals, a risk indicator that represents a. likelihood that the request may be illegitimate. The system can provide a responsive message to control reversal of the previously executed interaction and based on the risk indicator.
A method includes determining, using a trained machine-learning model, a risk indicator for a target entity from predictor variables associated with the target entity. The machine-learning model is trained based on a bi-partite similarity graph comprising a first set of nodes and a second set of nodes, where a first set of edges connects the first set of nodes and represents similarities between the nodes of the first set of nodes, a second set of edges connects the first set of nodes to the second set of nodes and represents relationships between the respective connected nodes, and a third set of edges connects the second set of nodes and represents similarities between the connected nodes of the second set of nodes. The method further includes generating and transmitting a responsive message comprising the risk indicator to control access of the target entity to a computing environment.
A system can be used to provide a responsive message for controlling an interaction dispute. The system can receive tokens from an optical character recognition model. The set of tokens can represent at least evidence data relating to an interaction dispute. The system can determine, using an artificial intelligence model, a first likelihood that represents a similarity between a subset of the tokens and the interaction dispute. The system can determine a second likelihood that traversing to the interaction dispute may result in success. The system can provide the responsive message that can control the interaction dispute based on the first likelihood and the second likelihood. The responsive message can include a response to the interaction dispute.
A computing system can generate and train a machine-learning model for risk assessment. The machine-learning model can be trained on semi-labelled graph data that may contain one or more isolated nodes generated from a tabularized data set. The computing system can use graph embeddings to compare pairs of nodes of the graph to determine a similarity between each pair of nodes. The similarity may be used to determine whether to create a synthetic edge between the pair of nodes. An additional hyperparameter may be used to tune the number of generated edges based on a desired graph density. The generated graph data may then be used to train a machine-learning model capable of generating a risk indicator for a target entity. Further, the risk indicators can be utilized to control the access by a target entity to an interactive computing environment for accessing services provided by one or more institutions.
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Computer software application for mobile devices, namely, computer software for monitoring, managing, and identifying changes in and risks to information relating to personal credit, fraud, and identity theft Providing temporary use of non-downloadable software for monitoring, managing, and identifying changes in and risks to information relating to personal credit, fraud, and identity theft
62.
ARTIFICIAL INTELLIGENCE MODEL FOR CONTROLLING INTERACTION DISPUTE
A system can be used to provide a responsive message for controlling an interaction dispute. The system can receive tokens from an optical character recognition model. The set of tokens can represent at least evidence data relating to an interaction dispute. The system can determine, using an artificial intelligence model, a first likelihood that represents a similarity between a subset of the tokens and the interaction dispute. The system can determine a second likelihood that traversing to the interaction dispute may result in success. The system can provide the responsive message that can control the interaction dispute based on the first likelihood and the second likelihood. The responsive message can include a response to the interaction dispute.
Techniques are described herein for applying natural language processing (NLP) techniques to time series data to derive attributes of an object for use with a machine-learning model. In one example, a system can receive a time series associated with an object over a time window, where the time series includes a set of discrete values. The system can then generate a time series encoding based on the time series. The system can provide the time series encoding as input to a trained natural language processing (NLP) model, which can generate one or more output embeddings based on the time series encoding. Next, the system can determine at least one attribute associated with the object based on the one or more output embeddings. The system can then provide the attributes for use with a machine-learning model, which may for example be configured to predict a future characteristic of the object.
A system can efficiently control access to an interactive computing environment. The system can receive authentication data of an authentication attempt associated with an entity. The system can determine, for the entity, a historical vector including features that include sub-features. The historical vector can be determined by generating synthetic data, generating weights, and determining probabilities. The synthetic data can be based on historical authentication attempts by entities other than the entity. The weights can correspond to sub-features of the historical vector. The probabilities can indicate a likelihood that a corresponding sub-feature is involved in the authentication attempt. The system can compare the historical vector to the authentication data. The system can generate a responsive message based on the comparison for controlling access to the interactive computing environment.
Systems and methods for automated historical risk assessment for risk mitigation in online access control are provided. An entity assessment server can receive a request to assess a risk indicator change from a first risk indicator to a second risk indicator. For each attribute used to generate the first risk indicator and second risk indicator, a first impact can be determined for changing from the first risk indicator to a third risk indicator between the first risk indicator and the second risk indicator. A second impact similarly can be determined for changing from the third risk indicator to the second risk indicator. Aggregating the first impact and the second impact can determine a total impact of each attribute. Assessment results can be generated to include a list of attributes ordered according to the respective total impact and transmitted to a remote computing device for use in improving the risk indicator.
G06F 21/57 - Certifying or maintaining trusted computer platforms, e.g. secure boots or power-downs, version controls, system software checks, secure updates or assessing vulnerabilities
66.
ENHANCED RANK-ORDER FOR RISK ASSESSMENT USING PARAMETERIZED DECAY
G06Q 10/04 - Forecasting or optimisation specially adapted for administrative or management purposes, e.g. linear programming or "cutting stock problem"
67.
ENHANCED RANK-ORDER FOR RISK ASSESSMENT USING PARAMETERIZED DECAY
A system can receive data about a target entity. The system can identify, for each data point included in the data, a set of parameters. The system can generate, based at least in part on the set of parameters, a rank-order list that includes the data. The rank-order list can be generated by a parameterized decay model. The system can provide the rank-order list to a user entity to control an interaction between the target entity and the user entity.
G06Q 10/04 - Forecasting or optimisation specially adapted for administrative or management purposes, e.g. linear programming or "cutting stock problem"
Search indexes can be automatically generated and used for expediting searching of a computerized database. For example, a system can access an inquiry dataset that includes relationships between prior inquiries and returned records from a database. The system can then generate a set of Boolean indexes based on the prior inquiries. The system can then identify frequent indexes that occur at least a threshold number of times in the set of Boolean indexes and that have estimated candidate sizes that are less than a threshold size. The system can then select the frequent index with the highest frequency from among the frequent indexes. The selected frequent index can be subsequently used to expedite searching of the database in response to receiving a search query associated with the frequent index from a client device.
In one example, a method includes building a first model including a first data set, the first data set excluding protected class. A second model is built including a second data set, the second data set including a subset of the first data set. The models may be used to generate and index score distributions. The distributions are compared to determine a self-report correlation. In response to determining the aggregate correlation is less than the first disparate impact threshold, the method generates a first prediction of disparate impact. The method then generates an aggregate correlation for attributes in the first model with the self-reported protected class attribute, the plurality of attributes being generated by the first model. In response to determining the aggregate correlation exceeds the second disparate impact threshold, the method generates a second prediction of disparate impact.
Various aspects of the present disclosure involve computing environments that provide third-party access-control support. For instance, an access-control computing system can access a secure identity repository having role history data from various contributor computing systems. The access-control computing system can compare an identified set of roles with a set of roles described by role history data for a target entity. The access-control computing system can determine, from the comparison, whether the target entity poses a security risk based on inconsistencies between the sets of roles, durations associated with the roles, or both. The access-control computing system can provide a client computing system with a dynamic access-control data structure that is generated based on the comparison. The dynamic access-control data structure allows the client computing system to output the security assessment to an end user or to otherwise facilitate further security measures with respect to the target entity.
H04L 41/0631 - Management of faults, events, alarms or notifications using root cause analysisManagement of faults, events, alarms or notifications using analysis of correlation between notifications, alarms or events based on decision criteria, e.g. hierarchy, tree or time analysis
72.
System and method for controlling access to a resource based on a compliance score
A system can generate one or more compliance graphs and a compliance score that can be used at least in risk assessment operations. The system can access a request to visualize target entity data and calculate a compliance score for the target entity. The system can access attribute data. The system can generate one or more compliance graphs and a compliance score using the attribute data. The system can compare the compliance score to a compliance threshold and transmit a message including the results of the comparison to a remote system for controlling access of the target entity to an interactive computing environment.
H04L 41/16 - Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks using machine learning or artificial intelligence
H04L 41/22 - Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks comprising specially adapted graphical user interfaces [GUI]
73.
DATA VALIDATION TECHNIQUES FOR SENSITIVE DATA MIGRATION ACROSS MULTIPLE PLATFORMS
Techniques for validating large amounts of sensitive data migrated across multiple platforms without revealing the content of the sensitive data are provided. For example, a processing device can transform data in a first data file stored on a first platform to common data formats. The processing device can generate a first set of hash values. The processing device can receive a second set of hash values for a second data file stored on a second platform. The processing device can compare the first set of hash values and the second set of hash values and cause the first data file or the second data file to be modified based on a difference between the sets of hash values.
H04L 9/06 - Arrangements for secret or secure communicationsNetwork security protocols the encryption apparatus using shift registers or memories for blockwise coding, e.g. D.E.S. systems
H04L 67/06 - Protocols specially adapted for file transfer, e.g. file transfer protocol [FTP]
H04L 67/1097 - Protocols in which an application is distributed across nodes in the network for distributed storage of data in networks, e.g. transport arrangements for network file system [NFS], storage area networks [SAN] or network attached storage [NAS]
74.
ARTIFICIAL INTELLIGENCE TECHNIQUES FOR IDENTIFYING IDENTITY MANIPULATION
A system can efficiently determine whether an identity is manipulated. The system can receive entity data and interaction data associated with a target entity. The system can determine, based on the entity data and the interaction data, one or more risk signals associated with the target entity using one or more artificial intelligence models. The system can generate a linked graph structure based on a first graph structure and a second graph structure each generated using the entity data and the interaction data. The system can apply the one or more risk signals to the linked graph structure to determine a risk indicator associated with the target entity. The system can provide a responsive message based on the risk indicator. The responsive message can be used to control access of the target entity to an interactive computing environment.
G06F 21/57 - Certifying or maintaining trusted computer platforms, e.g. secure boots or power-downs, version controls, system software checks, secure updates or assessing vulnerabilities
76.
ARTIFICIAL INTELLIGENCE TECHNIQUES FOR IDENTIFYING IDENTITY MANIPULATION
A system can efficiently determine whether an identity is manipulated. The system can receive entity data and interaction data associated with a target entity. The system can determine, based on the entity data and the interaction data, one or more risk signals associated with the target entity using one or more artificial intelligence models. The system can generate a linked graph structure based on a first graph structure and a second graph structure each generated using the entity data and the interaction data. The system can apply the one or more risk signals to the linked graph structure to determine a risk indicator associated with the target entity. The system can provide a responsive message based on the risk indicator. The responsive message can be used to control access of the target entity to an interactive computing environment.
A method can be used to predict risk and provide explainable outcomes using machine learning based on wavelet analysis. A risk prediction model can be applied to time-series data for an attribute associated with a target entity to generate a risk indicator for the target entity. The risk prediction model can include a feature learning model and a risk classification model configured to generate the risk indicator as output. Parameters of the feature learning model can be accessed and a plurality of basis functions of a wavelet transformation can be applied on the parameters of the feature learning model to generate a set of parameter wavelet coefficients. Explanatory data can be generated for the risk indicator based on the set of parameter wavelet coefficients. A responsive message can be transmitted to a remote computing device including the risk indicator and the explanatory data for use in controlling access of the target entity to an interactive computing environment.
Certain aspects involve building timing-prediction models for predicting timing of events that can impact one or more operations of machine-implemented environments. For instance, a computing system can generate program code executable by a host system for modifying host system operations based on the timing of a target event. The program code, when executed, can cause processing hardware to a compute set of probabilities for the target event by applying a set of trained timing-prediction models to predictor variable data. A time of the target event can be computed from the set of probabilities. To generate the program code, the computing system can build the set of timing-prediction models from training data. Building each timing-prediction model can include training the timing-prediction model to predict one or more target events for a different time bin within the training window. The computing system can generate and output program code implementing the models' functionality.
In some aspects, a computing system can use a machine learning model for resource management. For example, the system can receive a request for a set of steps associated with a target model output of a machine learning model. The request can include a starting input feature set and a number of steps. For each of the number of steps, the system can calculate a change to one or more features from the starting input feature set to arrive at the target model output based on a current position in feature space of the machine learning model. The system can update a feature vector by applying the change to the features of the starting input feature set and transmitting the set of steps. The system can then cause a resource of the external computing system to transition toward a position defined by the target model output.
36 - Financial, insurance and real estate services
42 - Scientific, technological and industrial services, research and design
45 - Legal and security services; personal services for individuals.
Goods & Services
(1) Consulting services in the field of life, health, property and casualty insurance; consulting services in the field of account origination and portfolio management; credit information services, namely, providing credit information relating to consumer or commercial applicants for credit, mortgage loans, utility services and employment and to fraud prevention; providing credit application processing; credit inquiry and consulting services; real estate appraisal services; providing an on-line credit information database featuring information relating to insurance, credit, financial data, mortgage loans, and employment, and to debt servicing; credit evaluation, analysis and alert services; and providing information relating to insurance, credit, mortgage loans and debt management; and excluding foreign exchange services and money market trading services.
(2) Providing non-downloadable software for risk analysis, identity verification, data analysis, risk modelling and analytics, decision making and analysis, in the field of fraud protection.
(3) Providing information over the Internet in the fields of identity verification and fraud protection; providing credit and financial fraud detection services.
82.
MACHINE-LEARNING TECHNIQUES FOR PREDICTING UNOBSERVABLE OUTPUTS
In some aspects, a computing system can generate and optimize a machine learning model to estimate an unobservable capacity of a target system or entity. The computing system can access training vectors which include training predictor variables, training performance indicators, and task quantities. A training performance indicator indicating performance outcome corresponding to the predictor variables and a task quantity associated with a task assigned to the target entity that leads to the training performance indicator. The machine learning model can be trained by performing adjustments of parameters of the machine learning model to minimize a loss function defined based on the training vectors. The trained machine learning model can be used to estimate the capacity of the target system or entity for handling tasks and be used in assigning tasks to the target entity according to the determined capacity.
In some aspects, a computing system can generate and optimize a machine learning model to estimate an unobservable capacity of a target system or entity. The computing system can access training vectors which include training predictor variables, training performance indicators, and task quantities. A training performance indicator indicating performance outcome corresponding to the predictor variables and a task quantity associated with a task assigned to the target entity that leads to the training performance indicator. The machine learning model can be trained by performing adjustments of parameters of the machine learning model to minimize a loss function defined based on the training vectors. The trained machine learning model can be used to estimate the capacity of the target system or entity for handling tasks and be used in assigning tasks to the target entity according to the determined capacity.
In some aspects, a machine learning (ML) model can be trained for risk assessment. The ML model can be trained to determine a risk indicator for a target entity from predictor variables associated with the target entity. The predictor variables are obtained from multiple sources with varying availability, and the training of the ML model is accomplished based on a multi-dimensional representation of common information from the set of data sources. Once generated, the risk indicator can be transmitted to a remote computing device in a responsive message for use in controlling access of the target entity to a computing environment.
In some aspects, a computing system can train a machine learning model for risk assessment. For example, the system can access a trained machine learning model to determine a final risk indicator of a target entity from a baseline data associated with the target entity and an alternative data associated with the target entity. The computing system can generate the final risk indicator of the target entity using the baseline data associated with the target entity and the alternative data associated with the target entity. The computing system can also transmit, to a remote computing device, a responsive message comprising at least the final risk indicator for use in controlling access of the target entity to one or more computing environments.
Various aspects involve explainable machine learning based on time-series transformation. For instance, a computing system accesses time-series data of a predictor variable associated with a target entity. The computing system generates a first set of transformed time-series data instances by applying a first family of transformations on the time-series data. Any non-negative linear combination of the first family of transformations forms an interpretable transformation of the time-series data. The computing system determines a risk indicator for the target entity indicating a level of risk associated with the target entity by inputting the first set of transformed time-series data instances into a machine learning model. The computing system transmits, to a remote computing device, a responsive message including the risk indicator. The risk indicator is usable for controlling access to one or more interactive computing environments by the target entity.
A system can efficiently control access to an interactive computing environment using similarity hashing. The system can receive a first similarity-preserving hash based on an interaction request associated with a target entity for requesting access to an interactive computing environment. The system can receive a second similarity-preserving hash based on entity data relating to an entity and including interaction data. The system can determine, based on a comparison between the first similarity-preserving hash and the second similarity-preserving hash, a likelihood that the interaction request is legitimate. The system can provide a responsive message based on a threshold relating to the likelihood, the responsive message usable to control access to the interactive computing environment.
H04L 9/32 - Arrangements for secret or secure communicationsNetwork security protocols including means for verifying the identity or authority of a user of the system
91.
TECHNIQUES FOR MITIGATING BACK PRESSURE, AUTO-SCALING THROUGHPUT, AND CONCURRENCY SCALING IN LARGE-SCALE AUTOMATED EVENT-DRIVEN DATA PIPELINES
Systems and methods for fine-tuned control over data transfer processes. An exemplary data transfer process may include: receiving a data stream at a storage service; receiving, at a first function, one or more notifications; in response to each notification, passing, by the first function, a message to a queue, the message comprising an address of a respective file within the storage service; receiving, at an invocation of a second function at a second computing service, one or more messages from the queue; retrieving, by the second function, data from one or more files based on the address in each of the one or more messages; and writing, by the second function, the data to a database. Systems and methods according to aspects of the present disclosure improve processes of transferring data from a data warehouse or database to a cloud-based database by mitigating back-pressure, auto-scaling throughput, and controlling concurrency scaling.
A system can generate an identity graph that can be arranged temporally for use at least in risk assessment operations and similar analyses. The system can receive a request to visualize entity data. The system can receive the entity data, which can include a set of identity data and a set of interaction data. The system can generate a temporal identity graph using the entity data. The temporal identity graph can temporally link the identity of the entity with the interactions associated with the entity. The system can generate a graphical user interface that may be configured to provide the temporal identity graph in response to the request to visualize the entity data. The graphical user interface can include interactive elements representing the temporal identity graph. Each interactive element can be selected by a user of the graphical user interface to display previously non-displayed information.
In some aspects, a record-matching computing system for detecting fragmented records is provided. The record-matching system is configured to identify a list of candidate records for merging from a set of data records. The record-matching system determines a matching decision for each pair of candidate records in the list and generates a graph. The graph includes nodes representing respective candidate records and edges connecting the nodes. Each edge represents a match between a pair of nodes connected by the edge according to the matching decisions. The record-matching system detects a connected component in the graph from which a qualified connected component is identified based on the minimum connectivity of the qualified connected component. The record-matching system updates the set of data records stored by merging candidate records represented by the nodes in the qualified connected component.
A system can generate a risk assessment associated with a target entity. The system can determine an attribute tier for each entity in a set of entities. For each attribute tier, the system can: generate a model configured to predict a percent change in the attribute over a time period for each entity in the respective attribute tier; determine the percent change in the attribute for each entity in the respective attribute tier using the model associated with the respective attribute tier; rank each entity in the attribute tier based on the predicted percent change in the attribute; and assign a score to each entity in the attribute tier based on the rank of the respective entity and on a preconfigured distribution. The system can determine a risk indicator based, in part, on the score.
In some aspects, a record-matching computing system for matching records to facilitate database search and fragmented records detection is provided. The record-matching computing system is configured to receiving a query record and search in a data repository storing data records for a record that matches the query record. The record-matching computing system retrieves a reference record from the data records and generates multiple identifier scores. Each identifier score measures a degree of matching between the corresponding identifiers in the query record and the reference record. The record-matching computing system generates an overall matching score by combining at least two of the identifier scores and determines the reference record as a match to the query record based on the overall matching score exceeding a threshold value.
In some aspects, a record-matching computing system for matching records to facilitate database search and fragmented records detection is provided. The record-matching computing system is configured to search for a data record that matches a query record. The record-matching computing system retrieves a reference record from data records and generates multiple identifier attributes for the query record and reference record, including identifier scores and compound scores. Each identifier score measures a degree of matching between the corresponding identifiers in the query record and reference record. A compound score is generated by combining two or more identifier scores. The record-matching computing system applies the identifier attributes to a machine learning model configured to predict a match classification based on input identifier attributes for a pair of data records. The record-matching server can identify the reference records as a match to the query record based on the match classification indicating a match.
In some aspects, techniques for creating representative and informative training datasets for the training of machine-learning models are provided. For example, a risk assessment system can receive a risk assessment query for a target entity. The risk assessment system can compute an output risk indicator for the target entity by applying a machine learning model to values of informative attributes associated with the target entity. The machine learning model may be trained using training samples selected from a representative and informative (RAI) dataset. The RAI dataset can be created by determining the informative attributes based on attributes used by a set of models and further extracting representative data records from an initial training dataset based on the determined informative attributes. The risk assessment system can transmit a responsive message including the output risk indicator for use in controlling access of the target entity to an interactive computing environment.
The present disclosure involves systems and methods for identity authentication across multiple institutions using a trusted mobile device as a proxy for a user login. In one example, the operations include identifying a request to trust a particular user associated with a first entity in a digital ID network. A set of personally identifiable information (PII) associated with the user is obtained via the first entity and an identity verification (IDV)/fraud risk analysis is performed. In response to satisfying the analysis, instructions are transmitted to the user to verify the identity via a mobile trust application on an associated mobile device. Upon verification, the mobile device is bound to the user within the digital ID network along with a digital ID associated with the particular user. The digital ID can be used by other entities registered within the digital ID network to authenticate the user.
G06F 21/62 - Protecting access to data via a platform, e.g. using keys or access control rules
G06K 7/10 - Methods or arrangements for sensing record carriers by electromagnetic radiation, e.g. optical sensingMethods or arrangements for sensing record carriers by corpuscular radiation
G06K 7/14 - Methods or arrangements for sensing record carriers by electromagnetic radiation, e.g. optical sensingMethods or arrangements for sensing record carriers by corpuscular radiation using light without selection of wavelength, e.g. sensing reflected white light
G06Q 20/32 - Payment architectures, schemes or protocols characterised by the use of specific devices using wireless devices
G06Q 20/40 - Authorisation, e.g. identification of payer or payee, verification of customer or shop credentialsReview and approval of payers, e.g. check of credit lines or negative lists
In some embodiments, a data visualization system accesses data entries associated with entities and obtained from multiple data sources. The data visualization system generates a classification map for the entities by classifying the entities into groups based on the data entries from the multiple data sources. The groups are arranged in the classification map according to values of the data entries of the entities. The data visualization system determines one or more metrics for the groups. The data visualization system further determines visualizations based on the classification map, each visualization representing a metric or the data entries from one of the data sources. The data visualization system generates, for inclusion in a user interface of the system, selectable interface elements configured for invoking an editing tool for updating the visualizations. The selectable interface elements for the visualizations are arranged in the respective visualizations according to the classification map.
Techniques are described herein for applying natural language processing (NLP) techniques to time series data to derive attributes of an object for use with a machine-learning model. In one example, a system can receive a time series associated with an object over a time window, where the time series includes a set of discrete values. The system can then generate a time series encoding based on the time series. The system can provide the time series encoding as input to a trained natural language processing (NLP) model, which can generate one or more output embeddings based on the time series encoding. Next, the system can determine at least one attribute associated with the object based on the one or more output embeddings. The system can then provide the attributes for use with a machine-learning model, which may for example be configured to predict a future characteristic of the object.