Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.
G10L 17/26 - Recognition of special voice characteristics, e.g. for use in lie detectorsRecognition of animal voices
G06F 21/32 - User authentication using biometric data, e.g. fingerprints, iris scans or voiceprints
G10L 15/02 - Feature extraction for speech recognitionSelection of recognition unit
G10L 17/00 - Speaker identification or verification techniques
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
G10L 17/08 - Use of distortion metrics or a particular distance between probe pattern and reference templates
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
G10L 25/60 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for measuring the quality of voice signals
H04M 3/22 - Arrangements for supervision, monitoring or testing
H04M 3/42 - Systems providing special services or facilities to subscribers
42 - Scientific, technological and industrial services, research and design
Goods & Services
Providing temporary use of non-downloadable cloud-based
software for authentication, biometric authentication,
liveness detection, behavioral analysis, risk scoring, and
fraud prevention in connection with software applications
and digital communications, namely, text, data, audio, and
video communications, telephony and voice communications,
contact center interactions, interactive voice response
(IVR) systems, conferencing, messaging, remote access,
screen sharing, data transfer, and document collaboration;
application service provider (ASP) services featuring
application programming interface (API) software for
authentication, biometric authentication, liveness
detection, behavioral analysis, risk scoring, and fraud
prevention in connection with software applications and
digital communications, including telephony and voice
communications, contact center interactions, and interactive
voice response (IVR) systems; software as a service (SaaS)
services featuring software for authentication, biometric
authentication, liveness detection, behavioral analysis,
risk scoring, identity verification, and fraud prevention in
connection with software applications and digital
communications, including telephony and voice
communications, contact center interactions, and interactive
voice response (IVR) systems; cloud computing featuring
software for detection and analysis of synthetic or
manipulated media, including audio, voice, music, and video
content generated or altered by artificial intelligence,
including deepfakes and replay attacks; providing temporary
use of online non-downloadable software for real-time
analysis of audio and video communications, including
telephony and voice communications and contact center
interactions, for detection of synthetic or manipulated
media and fraud; providing temporary use of online
non-downloadable software for user authentication, biometric
authentication, liveness detection, and identity and
geolocation verification in connection with telephony and
voice communications, contact center interactions, and
interactive voice response (IVR) systems; technical support
services, namely, remote and on-site installation,
deployment, implementation, management, and maintenance of
computer software systems for authentication, biometric
authentication, liveness detection, behavioral analysis,
risk scoring, and fraud prevention.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
42 - Scientific, technological and industrial services, research and design
Goods & Services
(1) Providing temporary use of non-downloadable cloud-based software for authentication, biometric authentication, liveness detection, behavioral analysis, risk scoring, and fraud prevention in connection with software applications and digital communications, namely, text, data, audio, and video communications, telephony and voice communications, contact center interactions, interactive voice response (IVR) systems, conferencing, messaging, remote access, screen sharing, data transfer, and document collaboration; application service provider (ASP) services featuring application programming interface (API) software for authentication, biometric authentication, liveness detection, behavioral analysis, risk scoring, and fraud prevention in connection with software applications and digital communications, including telephony and voice communications, contact center interactions, and interactive voice response (IVR) systems; software as a service (SaaS) services featuring software for authentication, biometric authentication, liveness detection, behavioral analysis, risk scoring, identity verification, and fraud prevention in connection with software applications and digital communications, including telephony and voice communications, contact center interactions, and interactive voice response (IVR) systems; cloud computing featuring software for detection and analysis of synthetic or manipulated media, including audio, voice, music, and video content generated or altered by artificial intelligence, including deepfakes and replay attacks; providing temporary use of online non-downloadable software for real-time analysis of audio and video communications, including telephony and voice communications and contact center interactions, for detection of synthetic or manipulated media and fraud; providing temporary use of online non-downloadable software for user authentication, biometric authentication, liveness detection, and identity and geolocation verification in connection with telephony and voice communications, contact center interactions, and interactive voice response (IVR) systems; technical support services, namely, remote and on-site installation, deployment, implementation, management, and maintenance of computer software systems for authentication, biometric authentication, liveness detection, behavioral analysis, risk scoring, and fraud prevention.
42 - Scientific, technological and industrial services, research and design
Goods & Services
Providing temporary use of non-downloadable cloud-based software for authentication, biometric authentication, liveness detection, behavioral analysis, risk scoring, and fraud prevention in connection with software applications and digital communications, namely, text, data, audio, and video communications, telephony and voice communications, contact center interactions, interactive voice response (IVR) systems, conferencing, messaging, remote access, screen sharing, data transfer, and document collaboration; Application service provider (ASP) services featuring application programming interface (API) software for authentication, biometric authentication, liveness detection, behavioral analysis, risk scoring, and fraud prevention in connection with software applications and digital communications, including telephony and voice communications, contact center interactions, and interactive voice response (IVR) systems; Software as a service (SaaS) services featuring software for authentication, biometric authentication, liveness detection, behavioral analysis, risk scoring, identity verification, and fraud prevention in connection with software applications and digital communications, including telephony and voice communications, contact center interactions, and interactive voice response (IVR) systems; Cloud computing featuring software for detection and analysis of synthetic or manipulated media, including audio, voice, music, and video content generated or altered by artificial intelligence, including deepfakes and replay attacks; Providing temporary use of online non-downloadable software for real-time analysis of audio and video communications, including telephony and voice communications and contact center interactions, for detection of synthetic or manipulated media and fraud; Providing temporary use of online non-downloadable software for user authentication, biometric authentication, liveness detection, and identity and geolocation verification in connection with telephony and voice communications, contact center interactions, and interactive voice response (IVR) systems; Technical support services, namely, remote and on-site installation, deployment, implementation, management, and maintenance of computer software systems for authentication, biometric authentication, liveness detection, behavioral analysis, risk scoring, and fraud prevention.
6.
SYSTEMS AND METHODS TO PREVENT DENIAL OF SERVICE ATTACKS FROM GENERATIVE AI VOICE BOTS
Embodiments disclosed herein include software processes and of machine-learning architectures for detecting and mitigating against synthetic speech instances. A computer analyzes audio speech data and metadata received with contact events associated with source identifiers. The computer executes machine-learning architecture(s) that determine whether the contact events likely include human-generated speech or machine-generated synthetic speech. The computer may determine the likelihood that contact events represent a DoS attack launched by a source device, by analyzing behavior features in metadata associated with the source identifier. The computer determines whether the contact events originated from the source user device having the source identifier launched a DoS attack and, if so, may update a blocklist. The blocklist may be stored in a database and includes one or more source identifiers that should be rejected or blocked at the current or inbound contact event or at future contact events for the particular source identifiers.
Embodiments disclosed herein include software processes and of machine-learning architectures for detecting and mitigating against synthetic speech instances. A computer analyzes audio speech data and metadata received with contact events associated with source identifiers. The computer executes machine-learning architecture(s) that determine whether the contact events likely include human-generated speech or machine-generated synthetic speech. The computer may determine the likelihood that contact events represent a DoS attack launched by a source device, by analyzing behavior features in metadata associated with the source identifier. The computer determines whether the contact events originated from the source user device having the source identifier launched a DoS attack and, if so, may update a blocklist. The blocklist may be stored in a database and includes one or more source identifiers that should be rejected or blocked at the current or inbound contact event or at future contact events for the particular source identifiers.
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
Disclosed are systems and methods including software processes executed by a server that detect machine-generated synthetic singing vocals in a vocal audio signal of an audio signal using a multi-stage machine-learning architecture. A singing detector identifies vocal segments containing singing. A singing liveness detector includes a fakeprint embedding extractor that extracts fakeprint feature vector embeddings representing artifacts of machine-generated vocal signals, scoring layers or classifier layers to generate a singing liveness score for identifying the likelihood a vocal signal is human-generated or synthetic. An optional singer detector includes a vocalprint embedding extractor that extracts vocalprint feature vector embeddings representing singer-specific vocal identity characteristics and generates a singer identification score or attribution score for identifying a particular singer in the vocal signal.
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
Disclosed are systems and methods including software processes executed by a server that detect machine-generated synthetic singing vocals in a vocal audio signal of an audio signal using a multi-stage machine-learning architecture. A singing detector identifies vocal segments containing singing. A singing liveness detector includes a fakeprint embedding extractor that extracts fakeprint feature vector embeddings representing artifacts of machine-generated vocal signals, scoring layers or classifier layers to generate a singing liveness score for identifying the likelihood a vocal signal is human-generated or synthetic. An optional singer detector includes a vocalprint embedding extractor that extracts vocalprint feature vector embeddings representing singer-specific vocal identity characteristics and generates a singer identification score or attribution score for identifying a particular singer in the vocal signal.
G10L 25/81 - Detection of presence or absence of voice signals for discriminating voice from music
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
Embodiments described herein provide for automatically authenticating telephone calls to an enterprise call center. The system disclosed herein builds on the trust of a data channel for the telephony channel. Certain types of authentication information can be received through the telephony channel, as well. But the mobile application associated with the call center system may provide additional or alternative forms of data through the data channel. The system may send requests to a mobile application of a device to provide information that can reliably be assumed to be coming from that particular device, such as a state of the device and/or a user's response to push notifications. In some cases, the authentication processes may be based on quantity and quality of matches between certain metadata or attributes expected to be received from a given device as compared to the metadata or attributes received.
Embodiments described herein provide for a voice biometrics system execute machine-learning architectures capable of passive, active, continuous, or static operations, or a combination thereof. Systems passively and/or continuously, in some cases in addition to actively and/or statically, enrolling speakers as the speakers speak into or around an edge device (e.g., car, television, radio, phone). The system identifies users on the fly without requiring a new speaker to mirror prompted utterances for reconfiguring operations. The system manages speaker profiles as speakers provide utterances to the system. Machine-learning architectures implement a passive and continuous voice biometrics system, possibly without knowledge of speaker identities. The system creates identities in an unsupervised manner, sometimes passively enrolling and recognizing known or unknown speakers. The system offers personalization and security across a wide range of applications, including media content for over-the-top services and IoT devices (e.g., personal assistants, vehicles), and call centers.
Embodiments include a computing device that executes software routines and/or one or more machine-learning architectures providing improved omni-channel authentication solutions. Embodiments include one or more computing devices that provide an authentication interface by which various communication channels may deposit contact or session data received via a first-channel session into a non-transitory storage medium of an authentication database for another channel to obtain and employ (e.g., verify users). This allows the customer to access an online data channel and enter the contact center through a telephony communication channel, but further allows the enterprise contact center systems to passively maintain access to various types of information about the user's identity captured from each contact channel, allowing the call center to request or capture authenticating information (e.g., voice biometrics) from both channels to employ authentication processes for one or both channels, such as voice biometrics authentication processes or other types of authentication functions.
Disclosed herein are embodiments of systems, methods, and products comprises an authentication server for caller ID verification. When a caller makes a phone call, the server receives the phone call and verifies whether the phone call is from a registered device associated with the phone number. The server queries the registered device to retrieve one or more current call states via an authentication function on the registered device. The server compares the states and/or state transitions to the observed states and/or state transitions of the phone call. If the registered device states and/or state transitions match the observed phone call states and/or state transitions, the server verifies that the phone call is from the registered device and not some imposter's device. If there is no such match, the server rejects the phone call before the call phone is connected or terminates the phone call after the phone call is connected.
H04M 3/436 - Arrangements for screening incoming calls
H04M 3/22 - Arrangements for supervision, monitoring or testing
H04M 3/42 - Systems providing special services or facilities to subscribers
H04M 19/04 - Current supply arrangements for telephone systems providing ringing current or supervisory tones, e.g. dialling tone or busy tone the ringing-current being generated at the substations
Disclosed are systems and methods including computing-processes executing machine-learning architectures for voice biometrics, in which the machine-learning architecture implements one or more language compensation functions. Embodiments include an embedding extraction engine (sometimes referred to as an “embedding extractor”) that extracts speaker embeddings and determines a speaker similarity score for determine or verifying the likelihood that speakers in different audio signals are the same speaker. The machine-learning architecture further includes a multi-class language classifier that determines a language likelihood score that indicates the likelihood that a particular audio signal includes a spoken language. The features and functions of the machine-learning architecture described herein may implement the various language compensation techniques to provide more accurate speaker recognition results, regardless of the language spoken by the speaker.
Disclosed are systems and methods including computing-processes, which may include layers of machine-learning architectures, for assessing risk for calls directed to call center systems using carrier signaling metadata. A computer evaluates carrier signaling metadata to perform various new risk-scoring techniques to determine riskiness of calls and authenticate calls. When determining a risk score for an incoming call is received at a call center system, the computer may obtain certain metadata values from inbound metadata, prior call metadata, or from third-party telecommunications services and executes processes for determining the risk score for the call. The risk score operations include several scoring components, including appliance print scoring, carrier detection scoring, ANI location detection scoring, location similarity scoring, and JIP-ANI location similarity scoring, among others.
Embodiments described herein provide for systems and methods for implementing a neural network architecture for spoof detection in audio signals. The neural network architecture contains a layers defining embedding extractors that extract embeddings from input audio signals. Spoofprint embeddings are generated for particular system enrollees to detect attempts to spoof the enrollee's voice. Optionally, voiceprint embeddings are generated for the system enrollees to recognize the enrollee's voice. The voiceprints are extracted using features related to the enrollee's voice. The spoofprints are extracted using features related to features of how the enrollee speaks and other artifacts. The spoofprints facilitate detection of efforts to fool voice biometrics using synthesized speech (e.g., deepfakes) that spoof and emulate the enrollee's voice.
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
G10L 17/08 - Use of distortion metrics or a particular distance between probe pattern and reference templates
G10L 17/00 - Speaker identification or verification techniques
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
18.
ONE TIME VOICE PASSPHRASE TO PROTECT AGAINST MAN-IN-THE-MIDDLE ATTACK
Embodiments described herein provide for automatically authenticating operation requests and end-users who submit operation requests during contact events. A server obtains an operation request for an operation originated at an end-user device. The server generates a voice-based one-time password (OTP) using contextual information associated with the requested operation. The server generates and transmits an OTP prompt having text representing the OTP for display at a user interface of the user device. The server receives a response including an audio signal that contains the recording of the user speaking the OTP text aloud. The server uses the audio signal to authenticate the user and the operation request based on the speaker's voice, the accuracy of the user speaking the OTP, and liveness or fraud detection features extracted from the audio signal or metadata from the user device.
G06Q 20/40 - Authorisation, e.g. identification of payer or payee, verification of customer or shop credentialsReview and approval of payers, e.g. check of credit lines or negative lists
G10L 17/04 - Training, enrolment or model building
19.
ONE TIME VOICE PASSPHRASE TO PROTECT AGAINST MAN-IN-THE-MIDDLE ATTACK
Embodiments described herein provide for automatically authenticating operation requests and end-users who submit operation requests during contact events. A server obtains an operation request for an operation originated at an end-user device. The server generates a voice-based one-time password (OTP) using contextual information associated with the requested operation. The server generates and transmits an OTP prompt having text representing the OTP for display at a user interface of the user device. The server receives a response including an audio signal that contains the recording of the user speaking the OTP text aloud. The server uses the audio signal to authenticate the user and the operation request based on the speaker's voice, the accuracy of the user speaking the OTP, and liveness or fraud detection features extracted from the audio signal or metadata from the user device.
Embodiments described herein provide for automatically authenticating operation requests and end-users who submit operation requests during contact events. A server obtains an operation request for an operation originated at an end-user device. The server generates a voice-based one-time password (OTP) using contextual information associated with the requested operation. The server generates and transmits an OTP prompt having text representing the OTP for display at a user interface of the user device. The server receives a response including an audio signal that contains the recording of the user speaking the OTP text aloud. The server uses the audio signal to authenticate the user and the operation request based on the speaker's voice, the accuracy of the user speaking the OTP, and liveness or fraud detection features extracted from the audio signal or metadata from the user device.
G06F 21/32 - User authentication using biometric data, e.g. fingerprints, iris scans or voiceprints
G10L 15/22 - Procedures used during a speech recognition process, e.g. man-machine dialog
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
21.
ONE TIME VOICE PASSPHRASE TO PROTECT AGAINST MAN-IN-THE-MIDDLE ATTACK
Embodiments described herein provide for automatically authenticating operation requests and end-users who submit operation requests during contact events. A server obtains an operation request for an operation originated at an end-user device. The server generates a voice-based one-time password (OTP) using contextual information associated with the requested operation. The server generates and transmits an OTP prompt having text representing the OTP for display at a user interface of the user device. The server receives a response including an audio signal that contains the recording of the user speaking the OTP text aloud. The server uses the audio signal to authenticate the user and the operation request based on the speaker's voice, the accuracy of the user speaking the OTP, and liveness or fraud detection features extracted from the audio signal or metadata from the user device.
Embodiments described herein provide for audio processing operations that evaluate characteristics of audio signals that are independent of the speaker's voice. A neural network architecture trains and applies discriminatory neural networks tasked with modeling and classifying speaker-independent characteristics. The task-specific models generate or extract feature vectors from input audio data based on the trained embedding extraction models. The embeddings from the task-specific models are concatenated to form a deep-phoneprint vector for the input audio signal. The DP vector is a low dimensional representation of the each of the speaker-independent characteristics of the audio signal and applied in various downstream operations.
H04L 9/14 - Arrangements for secret or secure communicationsNetwork security protocols using a plurality of keys or algorithms
H04L 9/32 - Arrangements for secret or secure communicationsNetwork security protocols including means for verifying the identity or authority of a user of the system
23.
ROBUST SPOOFING DETECTION SYSTEM USING DEEP RESIDUAL NEURAL NETWORKS
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 25/27 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique
25.
CHANNEL-COMPENSATED LOW-LEVEL FEATURES FOR SPEAKER RECOGNITION
A system for generating channel-compensated features of a speech signal includes a channel noise simulator that degrades the speech signal, a feed forward convolutional neural network (CNN) that generates channel-compensated features of the degraded speech signal, and a loss function that computes a difference between the channel-compensated features and handcrafted features for the same raw speech signal. Each loss result may be used to update connection weights of the CNN until a predetermined threshold loss is satisfied, and the CNN may be used as a front-end for a deep neural network (DNN) for speaker recognition/verification. The DNN may include convolutional layers, a bottleneck features layer, multiple fully-connected layers, and an output layer. The bottleneck features may be used to update connection weights of the convolutional layers, and dropout may be applied to the convolutional layers.
G10L 17/20 - Pattern transformations or operations aimed at increasing system robustness, e.g. against channel noise or different working conditions
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech ("deepfakes"). The server applies a machine-learning architecture that includes a segmentation engine trained to parse an audio signal into voiced- speech segments and unvoiced-speech segments. Each segment type is analyzed by respective deepfake detectors. A first deepfake detector generates a first risk score for the voiced-speech segment, while a second deepfake detector generates a second risk score for the unvoiced- speech segment. The machine-learning architecture includes fusion layers to algorithmically combine the risk scores to determine and overall risk score. In training, the server uses loss functions to calculate losses indicating distances or discrepancies between the generated risk scores and expected risk scores provided by training labels. Based on the loss, the server updates the parameters of the respective deepfake detectors or segmentation engine.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech ("deepfakes"). Embodiments implement a machine-learning architecture having a diffusion model that generates purified features that are fed to a deepfake detection model. The machine-learning architecture includes input layers that convert an audio signal into a Gaussian or frequency space representation (e.g., log spectrogram) to extract a set of initial features indicative of spoofing or deepfake attacks. The diffusion model identifies adversarial noise on the audio signal in the initial features and generates purified features or clean version of the input audio signal. A deepfake detector includes a neural network architecture and classifier programmed and trained to generate a deepfake detection score and classify the audio signal as genuine or fraudulent using the purified features.
G10L 19/018 - Audio watermarking, i.e. embedding inaudible data in the audio signal
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
G06F 21/82 - Protecting input, output or interconnection devices
Disclosed are systems and methods including software processes executed by a server that implement a machine-learning architecture for audio source tracing for deepfake detection. The computer extracts a feature vector representing features of the input audio signal. The machine-learning architecture includes one or more embedding extractors for extracting one or more feature vectors from the input audio signal. An attribute detector ingests an embedding and scoring layers generate a source-indicating attribute score. A source tracer includes a multi-class classifier to generate a signal source score using the attribute scores and generates a signal source class.
G10L 19/018 - Audio watermarking, i.e. embedding inaudible data in the audio signal
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
29.
DIFFUSION-BASED AUDIO PURIFICATION FOR DEFENDING AGAINST ADVERSARIAL DEEPFAKE ATTACKS
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”). Embodiments implement a machine-learning architecture having a diffusion model that generates purified features that are fed to a deepfake detection model. The machine-learning architecture includes input layers that convert an audio signal into a Gaussian or frequency space representation (e.g., log spectrogram) to extract a set of initial features indicative of spoofing or deepfake attacks. The diffusion model identifies adversarial noise on the audio signal in the initial features and generates purified features or clean version of the input audio signal. A deepfake detector includes a neural network architecture and classifier programmed and trained to generate a deepfake detection score and classify the audio signal as genuine or fraudulent using the purified features.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”). The server applies a machine-learning architecture that includes a segmentation engine trained to parse an audio signal into voiced-speech segments and unvoiced-speech segments. Each segment type is analyzed by respective deepfake detectors. A first deepfake detector generates a first risk score for the voiced-speech segment, while a second deepfake detector generates a second risk score for the unvoiced-speech segment. The machine-learning architecture includes fusion layers to algorithmically combine the risk scores to determine and overall risk score. In training, the server uses loss functions to calculate losses indicating distances or discrepancies between the generated risk scores and expected risk scores provided by training labels. Based on the loss, the server updates the parameters of the respective deepfake detectors or segmentation engine.
Disclosed are systems and methods including software processes executed by a server that implement a machine-learning architecture for audio source tracing for deepfake detection. The computer extracts a feature vector representing features of the input audio signal. The machine-learning architecture includes one or more embedding extractors for extracting one or more feature vectors from the input audio signal. An attribute detector ingests an embedding and scoring layers generate a source-indicating attribute score. A source tracer includes a multi-class classifier to generate a signal source score using the attribute scores and generates a signal source class.
Disclosed are systems and methods including software processes executed by a server that implement a machine-learning architecture for audio source tracing for deepfake detection. The computer extracts a feature vector representing features of the input audio signal. The machine-learning architecture includes one or more embedding extractors for extracting one or more feature vectors from the input audio signal. An attribute detector ingests an embedding and scoring layers generate a source-indicating attribute score. A source tracer includes a multi-class classifier to generate a signal source score using the attribute scores and generates a signal source class.
G10L 19/018 - Audio watermarking, i.e. embedding inaudible data in the audio signal
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
G06F 21/82 - Protecting input, output or interconnection devices
Embodiments described herein provide for a fraud detection engine for detecting various types of fraud at a call center and a fraud importance engine for tailoring the fraud detection operations to relative importance of fraud events. Fraud importance engine determines which fraud events are comparative more important than others. The fraud detection engine comprises machine-learning models that consume contact data and fraud importance information for various anti-fraud processes. The fraud importance engine calculates importance scores for fraud events based on user-customized attributes, such as fraud-type or fraud activity. The fraud importance scores are used in various processes, such as model training, model selection, and selecting weights or hyper-parameters for the ML models, among others. The fraud detection engine uses the importance scores to prioritize fraud alerts for review. The fraud importance engine receives detection feedback, which contacts involved false negatives, where fraud events were undetected but should have been detected.
Embodiments include a computing device that executes software routines and/or one or more machine-learning architectures including obtaining training audio signals having corresponding training impulse responses associated with reverberation degradation, training a machine-learning model of a presentation attack detection engine to generate one or more acoustic parameters by executing the presentation attack detection engine using the training impulse responses of the training audio signals and a loss function, obtaining an audio signal having an acoustic impulse response associated with reverberation degradation caused by one or more rooms, generating the one or more acoustic parameters for the audio signal by executing the machine-learning model using the audio signal as input, and generating an attack score for the audio signal based upon the one or more parameters generated by the machine learning model.
42 - Scientific, technological and industrial services, research and design
Goods & Services
Providing temporary use of non-downloadable cloud-based software for use in authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; application service provider featuring application programming interface (API) software for authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; software as a service (SaaS) services featuring software for authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; remote and on-site installation, deployment, implementation, management and maintenance of computer software systems for use in authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration
36.
END-TO-END SPEAKER RECOGNITION USING DEEP NEURAL NETWORK
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
42 - Scientific, technological and industrial services, research and design
Goods & Services
Providing temporary use of non-downloadable cloud-based software for use in authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; application service provider featuring application programming interface (API) software for authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; software as a service (SaaS) services featuring software for authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; remote and on-site installation, deployment, implementation, management and maintenance of computer software systems for use in authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration
42 - Scientific, technological and industrial services, research and design
Goods & Services
Providing temporary use of non-downloadable cloud-based software for use in authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; application service provider featuring application programming interface (API) software for authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; software as a service (SaaS) services featuring software for authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration; remote and on-site installation, deployment, implementation, management and maintenance of computer software systems for use in authentication and fraud prevention for software applications, text, data, audio, and video communications, interactions and transactions, video, audio, and data conferencing, instant messaging and discussion forums, remote computer access and screen sharing, data transfer, and document and data collaboration
39.
METHOD AND APPARATUS FOR THREAT IDENTIFICATION THROUGH ANALYSIS OF COMMUNICATIONS SIGNALING, EVENTS, AND PARTICIPANTS
Disclosed are systems and methods including software processes executed by a server for obtaining, by a computer, an audio signal including synthetic speech, extracting, by the computer, metadata from a watermark of the audio signal by applying a set of keys associated with a plurality of text-to-speech (TTS) services to the audio signal, the metadata indicating an origin of the synthetic speech in the audio signal, and generating, by the computer, based on the extracted metadata, a notification indicating that the audio signal includes the synthetic speech.
Embodiments disclosed herein include software processes executed by a computer for encoding and decoding watermarks for a speech signal in a call signal communicated via telephony channels. An encoder uses Linear Predictive Coding (LPC) to analyzes the call signal’s spectral envelope and embeds the watermark into the LPC log-spectrum of the speech signal of the call signal. The encoder may reduce the watermark’s strength at a formant peak of the speech signal, balancing the watermark’s robustness and detectability. A deep decoder includes a neural network architecture trained on watermarked and watermark-free speech signals having various types of degradation to extract a feature vector of a call signal and compute a watermark detection score for one or more frames or for the call signal. At inference time, the deep decoder detects the watermark when the watermark detection score satisfies a detection threshold.
Embodiments described herein provide for passive caller verification and/or passive fraud risk assessments for calls to customer call centers. Systems and methods may be used in real time as a call is coming into a call center. An analytics server of an analytics service looks at the purported Caller ID of the call, as well as the unaltered carrier metadata, which the analytics server then uses to generate or retrieve one or more probability scores using one or more lookup tables and/or a machine-learning model. A probability score indicates the likelihood that information derived using the Caller ID information has occurred or should occur given the carrier metadata received with the inbound call. The one or more probability scores be used to generate a risk score for the current call that indicates the probability of the call being valid (e.g., originated from a verified caller or calling device, non-fraudulent).
Embodiments described herein provide for a machine-learning architecture for modeling quality measures for enrollment signals. Modeling these enrollment signals enables the machine-learning architecture to identify deviations from expected or ideal enrollment signal in future test phase calls. These differences can be used to generate quality measures for the various audio descriptors or characteristics of audio signals. The quality measures can then be fused at the score-level with the speaker recognition's embedding comparisons for verifying the speaker. Fusing the quality measures with the similarity scoring essentially calibrates the speaker recognition's outputs based on the realities of what is actually expected for the enrolled caller and what was actually observed for the current inbound caller.
G10L 25/60 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for measuring the quality of voice signals
42 - Scientific, technological and industrial services, research and design
Goods & Services
Cloud computing featuring software for use in the detection
of various types of synthetic content such as voices, audio,
music, or video including deepfakes, machine-based replay,
and other media manipulated or generated by artificial
intelligence.
45.
ROBUST SPREAD-SPECTRUM SPEECH WATERMARKING USING LINEAR PREDICTION AND DEEP SPECTRAL SHAPING
Embodiments disclosed herein include software processes executed by a computer for encoding and decoding watermarks for a speech signal in a call signal communicated via telephony channels. An encoder uses Linear Predictive Coding (LPC) to analyzes the call signal's spectral envelope and embeds the watermark into the LPC log-spectrum of the speech signal of the call signal. The encoder may reduce the watermark's strength at a formant peak of the speech signal, balancing the watermark's robustness and detectability. A deep decoder includes a neural network architecture trained on watermarked and watermark-free speech signals having various types of degradation to extract a feature vector of a call signal and compute a watermark detection score for one or more frames or for the call signal. At inference time, the deep decoder detects the watermark when the watermark detection score satisfies a detection threshold.
G10L 19/018 - Audio watermarking, i.e. embedding inaudible data in the audio signal
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
46.
METHOD AND APPARATUS FOR THREAT IDENTIFICATION THROUGH ANALYSIS OF COMMUNICATIONS SIGNALING, EVENTS, AND PARTICIPANTS
Aspects of the invention determining a threat score of a call traversing a telecommunications network by leveraging the signaling used to originate, propagate and terminate the call. Outer-edge data utilized to originate the call may be analyzed against historical, or third party real-time data to determine the propensity of calls originating from those facilities to be categorized as a threat. Storing the outer edge data before the call is sent over the communications network permits such data to be preserved and not subjected to manipulations during traversal of the communications network. This allows identification of threat attempts based on the outer edge data from origination facilities, thereby allowing isolation of a compromised network facility that may or may not be known to be compromised by its respective network owner. Other aspects utilize inner edge data from an intermediate node of the communications network which may be analyzed against other inner edge data from other intermediate nodes and/or outer edge data.
42 - Scientific, technological and industrial services, research and design
Goods & Services
(1) Cloud computing featuring software for use in the detection of various types of synthetic content such as voices, audio, music, or video including deepfakes, machine-based replay, and other media manipulated or generated by artificial intelligence.
According to an embodiment of the disclosure, a toll-free telecommunications validation system determines a confidence value that an incoming phone call to an enterprises' toll-free number is originating from the station it purports to be by incorporating one or more layers of signals and data in determining said confidence value. The data and signals can include one or more call identifiers and/or toll-free call routing logs, service control point (SCP) signals and data, service data point (SDP) signals and data, dialed number information service (DNIS) signals and data, session initiation protocol (SIP) signals and data, carrier identification code (CIC) signals and data, location routing number (LRN) signals and data, jurisdiction information parameter (JIP) signals and data, charge number (CN) signals and data, billing number (BN) signals and data, and originating carrier information (such as information derived from the ANI).
The embodiments execute machine-learning architectures for biometric-based identity recognition (e.g., speaker recognition, facial recognition) and deepfake detection (e.g., speaker deepfake detection, facial deepfake detection). The machine-learning architecture includes layers defining multiple scoring components, including sub-architectures for speaker deepfake detection, speaker recognition, facial deepfake detection, facial recognition, and lip-sync estimation engine. The machine-learning architecture extracts and analyzes various types of low-level features from both audio data and visual data, combines the various scores, and uses the scores to determine the likelihood that the audiovisual data contains deepfake content and the likelihood that a claimed identity of a person in the video matches to the identity of an expected or enrolled person. This enables the machine-learning architecture to perform identity recognition and verification, and deepfake detection, in an integrated fashion, for both audio data and visual data.
The embodiments execute machine-learning architectures for biometric-based identity recognition (e.g., speaker recognition, facial recognition) and deepfake detection (e.g., speaker deepfake detection, facial deepfake detection). The machine-learning architecture includes layers defining multiple scoring components, including sub-architectures for speaker deepfake detection, speaker recognition, facial deepfake detection, facial recognition, and lip-sync estimation engine. The machine-learning architecture extracts and analyzes various types of low-level features from both audio data and visual data, combines the various scores, and uses the scores to determine the likelihood that the audiovisual data contains deepfake content and the likelihood that a claimed identity of a person in the video matches to the identity of an expected or enrolled person. This enables the machine-learning architecture to perform identity recognition and verification, and deepfake detection, in an integrated fashion, for both audio data and visual data.
Disclosed are systems and methods including software processes executed by a server for obtaining, by a computer, an audio signal including synthetic speech, extracting, by the computer, metadata from a watermark of the audio signal by applying a set of keys associated with a plurality of text-to-speech (TTS) services to the audio signal, the metadata indicating an origin of the synthetic speech in the audio signal, and generating, by the computer, based on the extracted metadata, a notification indicating that the audio signal includes the synthetic speech.
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.
G10L 17/26 - Recognition of special voice characteristics, e.g. for use in lie detectorsRecognition of animal voices
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
G10L 25/60 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for measuring the quality of voice signals
H04M 3/22 - Arrangements for supervision, monitoring or testing
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.
G10L 17/00 - Speaker identification or verification techniques
G06F 21/32 - User authentication using biometric data, e.g. fingerprints, iris scans or voiceprints
G10L 15/02 - Feature extraction for speech recognitionSelection of recognition unit
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
G10L 17/08 - Use of distortion metrics or a particular distance between probe pattern and reference templates
G10L 17/26 - Recognition of special voice characteristics, e.g. for use in lie detectorsRecognition of animal voices
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
G10L 25/60 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for measuring the quality of voice signals
H04M 3/22 - Arrangements for supervision, monitoring or testing
H04M 3/42 - Systems providing special services or facilities to subscribers
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.
G10L 17/26 - Recognition of special voice characteristics, e.g. for use in lie detectorsRecognition of animal voices
G06F 21/32 - User authentication using biometric data, e.g. fingerprints, iris scans or voiceprints
G10L 15/02 - Feature extraction for speech recognitionSelection of recognition unit
G10L 17/00 - Speaker identification or verification techniques
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
G10L 17/08 - Use of distortion metrics or a particular distance between probe pattern and reference templates
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
G10L 25/60 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for measuring the quality of voice signals
H04M 3/22 - Arrangements for supervision, monitoring or testing
H04M 3/42 - Systems providing special services or facilities to subscribers
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.
G10L 17/26 - Recognition of special voice characteristics, e.g. for use in lie detectorsRecognition of animal voices
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
G10L 17/08 - Use of distortion metrics or a particular distance between probe pattern and reference templates
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech ("deepfakes") in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech ("deepfakes") in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice "liveness" detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Systems, methods, and computer-readable media for call classification and for training a model for call classification, an example method comprising: receiving DTMF information from a plurality of calls; determining, for each of the calls, a feature vector including statistics based on DTMF information such as DTMF residual signal comprising channel noise and additive noise; training a model for classification; comparing a new call feature vector to the model; predicting a device type and geographic location based on the comparison of the new call feature vector to the model; classifying the call as spoofed or genuine; and authenticating a call or altering an IVR call flow.
H04M 3/22 - Arrangements for supervision, monitoring or testing
H04M 3/493 - Interactive information services, e.g. directory enquiries
H04M 7/12 - Arrangements for interconnection between switching centres for working between exchanges having different types of switching equipment, e.g. power-driven and step by step or decimal and non-decimal
H04M 15/06 - Recording class or number of calling party or called party
H04Q 1/45 - Signalling arrangementsManipulation of signalling currents using AC with voice-band signalling frequencies using multi-frequency signalling
H04Q 3/70 - Identification of class of calling subscriber
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
Embodiments include a computing device that executes software routines and/or one or more machine-learning architectures including obtaining training audio signals having corresponding training impulse responses associated with reverberation degradation, training a machine-learning model of a presentation attack detection engine to generate one or more acoustic parameters by executing the presentation attack detection engine using the training impulse responses of the training audio signals and a loss function, obtaining an audio signal having an acoustic impulse response associated with reverberation degradation caused by one or more rooms, generating the one or more acoustic parameters for the audio signal by executing the machine-learning model using the audio signal as input, and generating an attack score for the audio signal based upon the one or more parameters generated by the machine-learning model.
G10L 25/18 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
G10L 25/69 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for evaluating synthetic or decoded voice signals
71.
Joint estimation of acoustic parameters from single-microphone speech
Embodiments described herein provide for end-to-end joint determination of degradation parameter scores for certain types of degradation. Degradation parameters include degradation describing additive noise and multiplicative noise such as Signal-to-Noise Ratio (SNR), reverberation time (T60), and Direct-to-Reverberant Ratio (DRR). Various neural network architectures are described such that the inherent interplay between the degradation parameters is considered in both the degradation parameter score and degradation score determination. The neural network architectures are trained according to computer generated audio datasets.
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
42 - Scientific, technological and industrial services, research and design
Goods & Services
Cloud computing featuring software for use in the detection of various types of synthetic content such as voices, audio, music, or video including deepfakes, machine-based replay, and other media manipulated or generated by artificial intelligence.
73.
SYSTEMS AND METHODS FOR CALL FRAUD ANALYSIS USING A MACHINE-LEARNING ARCHITECTURE AND MAINTAINING CALLER ANI PRIVACY
Disclosed are systems and methods including processes executed by a server that executes software routines for machine-learning architectures that receive call-invite messages containing data from a terminating carrier. The server a caller ANI and types of call data. The server further requests data from a telephony database. The server applies and executes the software programming of the machine-learning architecture on the call data (from the terminating carrier) and the portability data (from the telephony database) to generate risk scores. The server stores the data and the risk scores into a request database, until a provider server requests the risk scores in a threat assessment request. The server returns a threat assessment message to the provider server in response to the threat assessment request. The threat assessment message includes information about the caller or caller device, and the risk scores, but not the caller ANI.
The present invention is directed to a deep neural network (DNN) having a triplet network architecture, which is suitable to perform speaker recognition. In particular, the DNN includes three feed-forward neural networks, which are trained according to a batch process utilizing a cohort set of negative training samples. After each batch of training samples is processed, the DNN may be trained according to a loss function, e.g., utilizing a cosine measure of similarity between respective samples, along with positive and negative margins, to provide a robust representation of voiceprints.
G10L 15/16 - Speech classification or search using artificial neural networks
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
Embodiments described herein provide for audio processing operations that evaluate characteristics of audio signals that are independent of the speaker's voice. A neural network architecture trains and applies discriminatory neural networks tasked with modeling and classifying speaker-independent characteristics. The task-specific models generate or extract feature vectors from input audio data based on the trained embedding extraction models. The embeddings from the task-specific models are concatenated to form a deep-phoneprint vector for the input audio signal. The DP vector is a low dimensional representation of the each of the speaker-independent characteristics of the audio signal and applied in various downstream operations.
Embodiments described herein provide for a fraud detection engine for detecting various types of fraud at a call center and a fraud importance engine for tailoring the fraud detection operations to relative importance of fraud events. Fraud importance engine determines which fraud events are comparative more important than others. The fraud detection engine comprises machine-learning models that consume contact data and fraud importance information for various anti-fraud processes. The fraud importance engine calculates importance scores for fraud events based on user-customized attributes, such as fraud-type or fraud activity. The fraud importance scores are used in various processes, such as model training, model selection, and selecting weights or hyper-parameters for the ML models, among others. The fraud detection engine uses the importance scores to prioritize fraud alerts for review. The fraud importance engine receives detection feedback, which contacts involved false negatives, where fraud events were undetected but should have been detected.
Utterances of at least two speakers in a speech signal may be distinguished and the associated speaker identified by use of diarization together with automatic speech recognition of identifying words and phrases commonly in the speech signal. The diarization process clusters turns of the conversation while recognized special form phrases and entity names identify the speakers. A trained probabilistic model deduces which entity name(s) correspond to the clusters.
Embodiments include a computing device that executes software routines and/or one or more machine-learning architectures including a neural network-based embedding extraction system that to produce an embedding vector representing a user's behavior's keypresses, where the system extracts the behaviorprint embedding vector using the keypress features that the system references later for authenticating users. Embodiments may extract and evaluate keypress features, such as keypress sequences, keypress pressure or volume, and temporal keypress features, such as the duration of keypresses and the interval between keypresses, among others. Some embodiments employ a deep neural network architecture that generates a behaviorprint embedding vector representation of the keypress duration and interval features that is used for enrollment and at inference time to authenticate users.
Embodiments include a computing device that executes software routines and/or one or more machine-learning architectures including a neural network-based embedding extraction system that to produce an embedding vector representing a user's behavior's keypresses, where the system extracts the behaviorprint embedding vector using the keypress features that the system references later for authenticating users. Embodiments may extract and evaluate keypress features, such as keypress sequences, keypress pressure or volume, and temporal keypress features, such as the duration of keypresses and the interval between keypresses, among others. Some embodiments employ a deep neural network architecture that generates a behaviorprint embedding vector representation of the keypress duration and interval features that is used for enrollment and at inference time to authenticate users.
Embodiments described herein provide for passive caller verification and/or passive fraud risk assessments for calls to customer call centers. Systems and methods may be used in real time as a call is coming into a call center. An analytics server of an analytics service looks at the purported Caller ID of the call, as well as the unaltered carrier metadata, which the analytics server then uses to generate or retrieve one or more probability scores using one or more lookup tables and/or a machine-learning model. A probability score indicates the likelihood that information derived using the Caller ID information has occurred or should occur given the carrier metadata received with the inbound call. The one or more probability scores be used to generate a risk score for the current call that indicates the probability of the call being valid (e.g., originated from a verified caller or calling device, non-fraudulent).
Embodiments described herein provide for systems and methods for implementing a neural network architecture for spoof detection in audio signals. The neural network architecture contains a layers defining embedding extractors that extract embeddings from input audio signals. Spoofprint embeddings are generated for particular system enrollees to detect attempts to spoof the enrollee's voice. Optionally, voiceprint embeddings are generated for the system enrollees to recognize the enrollee's voice. The voiceprints are extracted using features related to the enrollee's voice. The spoofprints are extracted using features related to features of how the enrollee speaks and other artifacts. The spoofprints facilitate detection of efforts to fool voice biometrics using synthesized speech (e.g., deepfakes) that spoof and emulate the enrollee's voice.
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
G10L 17/08 - Use of distortion metrics or a particular distance between probe pattern and reference templates
Embodiments described herein provide for a computer that detects one or more keywords of interest using acoustic features, to detect or query commonalities across multiple fraud calls. Embodiments described herein may implement unsupervised keyword spotting (UKWS) or unsupervised word discovery (UWD) in order to identify commonalities across a set of calls, where both UKWS and UWD employ Gaussian Mixture Models (GMM) and one or more dynamic time-warping algorithms. A user may indicate a training exemplar or occurrence of call-specific information, referred to herein as “a named entity,” such as a person's name, an account number, account balance, or order number. The computer may perform a redaction process that computationally nullifies the import of the named entity in the modeling processes described herein.
Embodiments include a computing device that executes software routines and/or one or more machine-learning architectures providing improved omni-channel authentication solutions. Embodiments include one or more computing devices that provide an authentication interface by which various communication channels may deposit contact or session data received via a first-channel session into a non-transitory storage medium of an authentication database for another channel to obtain and employ (e.g., verify users). This allows the customer to access an online data channel and enter the contact center through a telephony communication channel, but further allows the enterprise contact center systems to passively maintain access to various types of information about the user's identity captured from each contact channel, allowing the call center to request or capture authenticating information (e.g., voice biometrics) from both channels to employ authentication processes for one or both channels, such as voice biometrics authentication processes or other types of authentication functions.
Embodiments include a computing device that executes software routines and/or one or more machine-learning architectures providing improved omni-channel authentication solutions. Embodiments include one or more computing devices that provide an authentication interface by which various communication channels may deposit contact or session data received via a first-channel session into a non-transitory storage medium of an authentication database for another channel to obtain and employ (e.g., verify users). This allows the customer to access an online data channel and enter the contact center through a telephony communication channel, but further allows the enterprise contact center systems to passively maintain access to various types of information about the user's identity captured from each contact channel, allowing the call center to request or capture authenticating information (e.g., voice biometrics) from both channels to employ authentication processes for one or both channels, such as voice biometrics authentication processes or other types of authentication functions.
G06Q 30/015 - Providing customer assistance, e.g. assisting a customer within a business location or via helpdesk
G10L 17/08 - Use of distortion metrics or a particular distance between probe pattern and reference templates
H04L 9/32 - Arrangements for secret or secure communicationsNetwork security protocols including means for verifying the identity or authority of a user of the system
85.
Carrier signaling based authentication and fraud detection
Disclosed are systems and methods including computing-processes, which may include layers of machine-learning architectures, for assessing risk for calls directed to call center systems using carrier signaling metadata. A computer evaluates carrier signaling metadata to perform various new risk-scoring techniques to determine riskiness of calls and authenticate calls. When determining a risk score for an incoming call is received at a call center system, the computer may obtain certain metadata values from inbound metadata, prior call metadata, or from third-party telecommunications services and executes processes for determining the risk score for the call. The risk score operations include several scoring components, including appliance print scoring, carrier detection scoring, ANI location detection scoring, location similarity scoring, and JIP-ANI location similarity scoring, among others.
Disclosed are systems and methods including computing-processes, which may include layers of machine-learning architectures, for assessing risk for calls directed to call center systems using carrier signaling metadata. A computer evaluates carrier signaling metadata to perform various new risk-scoring techniques to determine riskiness of calls and authenticate calls. When determining a risk score for an incoming call is received at a call center system, the computer may obtain certain metadata values from inbound metadata, prior call metadata, or from third-party telecommunications services and executes processes for determining the risk score for the call. The risk score operations include several scoring components, including appliance print scoring, carrier detection scoring, ANI location detection scoring, location similarity scoring, and JIP-ANI location similarity scoring, among others.
Utterances of at least two speakers in a speech signal may be distinguished and the associated speaker identified by use of diarization together with automatic speech recognition of identifying words and phrases commonly in the speech signal. The diarization process clusters turns of the conversation while recognized special form phrases and entity names identify the speakers. A trained probabilistic model deduces which entity name(s) correspond to the clusters.
A system for generating channel-compensated features of a speech signal includes a channel noise simulator that degrades the speech signal, a feed forward convolutional neural network (CNN) that generates channel-compensated features of the degraded speech signal, and a loss function that computes a difference between the channel-compensated features and handcrafted features for the same raw speech signal. Each loss result may be used to update connection weights of the CNN until a predetermined threshold loss is satisfied, and the CNN may be used as a front-end for a deep neural network (DNN) for speaker recognition/verification. The DNN may include convolutional layers, a bottleneck features layer, multiple fully-connected layers, and an output layer. The bottleneck features may be used to update connection weights of the convolutional layers, and dropout may be applied to the convolutional layers.
G10L 17/20 - Pattern transformations or operations aimed at increasing system robustness, e.g. against channel noise or different working conditions
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 17/04 - Training, enrolment or model building
A method of obtaining and automatically providing secure authentication information includes registering a client device over a data line, storing information and a changeable value for authentication in subsequent telephone-only transactions. In the subsequent transactions, a telephone call placed from the client device to an interactive voice response server is intercepted and modified to include dialing of a delay and at least a passcode, the passcode being based on the unique information and the changeable value, where the changeable value is updated for every call session. The interactive voice response server forwards the passcode and a client device identifier to an authentication function, which compares the received passcode to plural passcodes generated based on information and iterations of a value stored in correspondence with the client device identifier. Authentication is confirmed when a generated passcode matches the passcode from the client device.
Embodiments described herein provide for evaluating call metadata and certificates of inbound calls for authentication. The computer identifies a service provider indicated by the SPID and/or the ANI (or other identifier) of the metadata and identifies a service provider indicated by the SPID and/or ANI (or other identifier) of the certificate, then compares identities of the service providers and/or compares the data values associated with the service providers (e.g., SPIDs, ANIs). Based on this comparison, the computer determines whether the service provider that signed the certificate is first-party signer (e.g., carrier) for the ANI or a third-party signer that is signing certificates as the first-party signer for the ANI.
Embodiments described herein provide for evaluating call metadata and certificates of inbound calls for authentication. The computer identifies a service provider indicated by the SPID and/or the ANI (or other identifier) of the metadata and identifies a service provider indicated by the SPID and/or ANI (or other identifier) of the certificate, then compares identities of the service providers and/or compares the data values associated with the service providers (e.g., SPIDs, ANIs). Based on this comparison, the computer determines whether the service provider that signed the certificate is first-party signer (e.g., carrier) for the ANI or a third-party signer that is signing certificates as the first-party signer for the ANI.
Embodiments described herein provide for systems and methods for verifying authentic JIPs associated with ANIs using CLLIs known to be associated with the ANIs, allowing a computer to authenticate calls using the verified JIPs, among various factors. The computer builds a trust model for JIPs by correlating unique CLLIs to JIPs. A malicious actor might spoof numerous ANIs mapped to a single CLLI, but the malicious actor is unlikely to spoof multiple CLLIs due to the complexity of spoofing the volumes of ANIs associated with multiple CLLIs, so the CLLIs can be trusted when determining whether a JIP is authentic. The computer identifies an authentic JIP when the trust model indicates that a number of CLLIs associated with the JIP satisfies one or more thresholds. A machine-learning architecture references the fact that the JIP is authentic as an authentication factor for downstream call authentication functions.
Embodiments described herein provide for systems and methods for verifying authentic JIPs associated with ANIs using CLLIs known to be associated with the ANIs, allowing a computer to authenticate calls using the verified JIPs, among various factors. The computer builds a trust model for JIPs by correlating unique CLLIs to JIPs. A malicious actor might spoof numerous ANIs mapped to a single CLLI, but the malicious actor is unlikely to spoof multiple CLLIs due to the complexity of spoofing the volumes of ANIs associated with multiple CLLIs, so the CLLIs can be trusted when determining whether a JIP is authentic. The computer identifies an authentic JIP when the trust model indicates that a number of CLLIs associated with the JIP satisfies one or more thresholds. A machine-learning architecture references the fact that the JIP is authentic as an authentication factor for downstream call authentication functions.
Embodiments described herein provide for performing a risk assessment using graph-derived features of a user interaction. A computer receives interaction information and infers information from the interaction based on information provided to the computer by a communication channel used in transmitting the interaction information. The computer may determine a claimed identity of the user associated with the user interaction. The computer may extract features from the inferred identity and claimed identity. The computer generates a graph representing the structural relationship between the communication channels and claimed identities associated with the inferred identity and claimed identity. The computer may extract additional features from the inferred identity and claimed identity using the graph. The computer may apply the features to a machine learning model to generate a risk score indicating the probability of a fraudulent interaction associated with the user interaction.
According to an embodiment of the disclosure, a toll-free telecommunications validation system determines a confidence value that an incoming phone call to an enterprises' toll-free number is originating from the station it purports to be, i.e., is not a spoofed call by incorporating one or more layers of signals and data in determining said confidence value, the data and signals including, but not limited to, toll-free call routing logs, service control point (SCP) signals and data, service data point (SDP) signals and data, dialed number information service (DNIS) signals and data, automatic number identification (ANI) signals and data, session initiation protocol (SIP) signals and data, carrier identification code (CIC) signals and data, location routing number (LRN) signals and data, jurisdiction information parameter (JIP) signals and data, charge number (CN) signals and data, billing number (BN) signals and data, and originating carrier information (such as information derived from the ANI, including, but not limited to, alternative service provider ID (ALTSPID), service provider ID (SPID), or operating company number (OCN)). In certain configurations said enterprise provides an ANI and DNIS associated with said incoming toll-free call, which is used to query a commercial toll-free telecommunications routing platform for any corresponding log entries. The existence of any such log entries, along with the originating carrier information in the event log entries do exist, is used to determine a confidence value that said incoming toll-free call is originating from the station it purports to be. As a result, said entities or enterprises operating a toll-free number may be provided a confidence value regarding an incoming telephone call, and using that confidence value, further determine whether or not to accept the authenticity of the incoming telephone call and/or based on said confidence value, service the incoming call differently.
Disclosed are systems and methods including computing-processes executing machine-learning architectures for voice biometrics, in which the machine-learning architecture implements one or more language compensation functions. Embodiments include an embedding extraction engine (sometimes referred to as an "embedding extractor") that extracts speaker embeddings and determines a speaker similarity score for determine or verifying the likelihood that speakers in different audio signals are the same speaker. The machine-learning architecture further includes a multi-class language classifier that determines a language likelihood score that indicates the likelihood that a particular audio signal includes a spoken language. The features and functions of the machine-learning architecture described herein may implement the various language compensation techniques to provide more accurate speaker recognition results, regardless of the language spoken by the speaker.
Disclosed are systems and methods including computing-processes executing machine-learning architectures for voice biometrics, in which the machine-learning architecture implements one or more language compensation functions. Embodiments include an embedding extraction engine (sometimes referred to as an "embedding extractor") that extracts speaker embeddings and determines a speaker similarity score for determine or verifying the likelihood that speakers in different audio signals are the same speaker. The machine-learning architecture further includes a multi-class language classifier that determines a language likelihood score that indicates the likelihood that a particular audio signal includes a spoken language. The features and functions of the machine-learning architecture described herein may implement the various language compensation techniques to provide more accurate speaker recognition results, regardless of the language spoken by the speaker.
G10L 25/60 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for measuring the quality of voice signals
Disclosed are systems and methods including computing-processes executing machine-learning architectures for voice biometrics, in which the machine-learning architecture implements one or more language compensation functions. Embodiments include an embedding extraction engine (sometimes referred to as an “embedding extractor”) that extracts speaker embeddings and determines a speaker similarity score for determine or verifying the likelihood that speakers in different audio signals are the same speaker. The machine-learning architecture further includes a multi-class language classifier that determines a language likelihood score that indicates the likelihood that a particular audio signal includes a spoken language. The features and functions of the machine-learning architecture described herein may implement the various language compensation techniques to provide more accurate speaker recognition results, regardless of the language spoken by the speaker.
Disclosed are systems and methods including computing-processes executing machine- learning architectures implementing label distribution loss functions to improve age estimation performance and generalization. The machine-learning architecture includes a front-end neural network architecture defining a speaker embedding extraction engine of the machine-learning architecture, and a backend neural network architecture defining an age estimation engine of the machine-learning architecture. The embedding extractor is trained to extract low-level acoustic features of a speaker's speech, such as mel-frequency cepstral coefficients (MFCCs), from audio signals, and then extract a feature vector or speaker embedding vector that mathematically represents the low-level features of the speaker. The age estimator is trained to generate an estimated age for the speaker and a Gaussian probability distribution around the estimated age, by applying the various types of layers of the age estimator on the speaker embedding.
Disclosed are systems and methods including computing-processes executing machine- learning architectures implementing label distribution loss functions to improve age estimation performance and generalization. The machine-learning architecture includes a front-end neural network architecture defining a speaker embedding extraction engine of the machine-learning architecture, and a backend neural network architecture defining an age estimation engine of the machine-learning architecture. The embedding extractor is trained to extract low-level acoustic features of a speaker's speech, such as mel-frequency cepstral coefficients (MFCCs), from audio signals, and then extract a feature vector or speaker embedding vector that mathematically represents the low-level features of the speaker. The age estimator is trained to generate an estimated age for the speaker and a Gaussian probability distribution around the estimated age, by applying the various types of layers of the age estimator on the speaker embedding.