09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable and recorded computer software platforms using
artificial intelligence for making podcasts and other
multimedia content; downloadable and recorded computer
software platforms using artificial intelligence for
recording, transcribing, editing, and mixing audio, video,
text, and other media content; downloadable and recorded
audio word processing computer software platforms using
artificial intelligence for enabling editing of sound files
and lyrics in text form; downloadable computer software
using artificial intelligence for use in prompting
questions, offering observations, and providing challenges
to users in the field of audio, video, text, and other media
content; downloadable computer software using artificial
intelligence for recording audio and transforming that audio
into text, audio, images, video, or combinations thereof;
downloadable digital image files of avatars; downloadable
computer software using artificial intelligence for creating
avatars. Providing on-line non-downloadable computer software using
artificial intelligence for making podcasts and other
multimedia content; platform as a service (PaaS) featuring
computer software platforms using artificial intelligence
for making podcasts and other multimedia content; providing
on-line non-downloadable computer software using artificial
intelligence for recording, transcribing, editing, and
mixing audio, video, text, and other media content; platform
as a service (PaaS) featuring computer software platforms
using artificial intelligence (AI) for recording,
transcribing, editing, and mixing audio, video, text, and
other media content; technical support services relating to
recording, transcribing, editing, and mixing audio, video,
text, and other media content, namely, troubleshooting in
the nature of diagnosing computer software problems using
artificial intelligence; providing on-line non-downloadable
computer software using artificial intelligence for enabling
editing of sound files and lyrics in text form; platform as
a service (PaaS) featuring computer software and audio and
word processing platforms using artificial intelligence for
enabling editing of sound files and lyrics in text form;
providing on-line non-downloadable computer software using
artificial intelligence for use in prompting questions,
offering observations, and providing challenges to users in
the field of audio, video, text, and other media content;
platform as a service (PaaS) featuring computer software and
audio and word processing platforms using artificial
intelligence for use in prompting questions, offering
observations, and providing challenges to users in the field
of audio, video, text, and other media content; providing
on-line non-downloadable computer software using artificial
intelligence for recording audio and transforming that audio
into text, audio, images, video, or combinations thereof;
providing on-line non-downloadable computer software using
artificial intelligence (AI) for creating avatars; platform
as a service (PaaS) featuring computer software and audio
and word processing platforms using artificial intelligence
for creating avatars.
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable and recorded computer software platforms using
artificial intelligence for making podcasts and other
multimedia content; downloadable and recorded computer
software platforms using artificial intelligence for
recording, transcribing, editing, and mixing audio, video,
text, and other media content; downloadable and recorded
audio word processing computer software platforms using
artificial intelligence for enabling editing of sound files
and lyrics in text form; downloadable computer software
using artificial intelligence for use in prompting
questions, offering observations, and providing challenges
to users in the field of audio, video, text, and other media
content; downloadable computer software using artificial
intelligence for recording audio and transforming that audio
into text, audio, images, video, or combinations thereof;
downloadable digital image files of avatars; downloadable
computer software using artificial intelligence for creating
avatars. Providing on-line non-downloadable computer software using
artificial intelligence for making podcasts and other
multimedia content; platform as a service (PaaS) featuring
computer software platforms using artificial intelligence
for making podcasts and other multimedia content; providing
on-line non-downloadable computer software using artificial
intelligence for recording, transcribing, editing, and
mixing audio, video, text, and other media content; platform
as a service (PaaS) featuring computer software platforms
using artificial intelligence (AI) for recording,
transcribing, editing, and mixing audio, video, text, and
other media content; technical support services relating to
recording, transcribing, editing, and mixing audio, video,
text, and other media content, namely, troubleshooting in
the nature of diagnosing computer software problems using
artificial intelligence; providing on-line non-downloadable
computer software using artificial intelligence for enabling
editing of sound files and lyrics in text form; platform as
a service (PaaS) featuring computer software and audio and
word processing platforms using artificial intelligence for
enabling editing of sound files and lyrics in text form;
providing on-line non-downloadable computer software using
artificial intelligence for use in prompting questions,
offering observations, and providing challenges to users in
the field of audio, video, text, and other media content;
platform as a service (PaaS) featuring computer software and
audio and word processing platforms using artificial
intelligence for use in prompting questions, offering
observations, and providing challenges to users in the field
of audio, video, text, and other media content; providing
on-line non-downloadable computer software using artificial
intelligence for recording audio and transforming that audio
into text, audio, images, video, or combinations thereof;
providing on-line non-downloadable computer software using
artificial intelligence (AI) for creating avatars; platform
as a service (PaaS) featuring computer software and audio
and word processing platforms using artificial intelligence
for creating avatars.
3.
APPROACHES TO MULTIMEDIA EDITING USING AN ARTIFICIAL INTELLIGENCE MODEL AND SYSTEMS FOR ACCOMPLISHING THE SAME
The disclosed technology uses a media production platform to edit multimedia files with an AI model (e.g., a neural network). The technology can remove retakes, identify highlight clips, and/or generate layouts for multimedia files. The technology can process audio transcripts to exclude retakes by generating a refined transcript and highlighting removed segments. Additionally, the technology can edit audiovisual files by generating scenes based on content and mapping the scenes to relevant layouts, dynamically adjusting based on user input. The technology can generate highlights by applying AI models to create clips and identify topics within the audiovisual file, producing an edited file indicative of the topics. The results, such as the edited files, are presented on the client device.
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
(1) Downloadable and recorded computer software platforms using artificial intelligence for making podcasts and other multimedia content; downloadable and recorded computer software platforms using artificial intelligence for recording, transcribing, editing, and mixing audio, video, text, and other media content; downloadable and recorded audio word processing computer software platforms using artificial intelligence for enabling editing of sound files and lyrics in text form; downloadable computer software using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; downloadable computer software using artificial intelligence for recording audio and transforming that audio into text, audio, images, video, or combinations thereof; downloadable digital image files of avatars; downloadable computer software using artificial intelligence for creating avatars. (1) Providing on-line non-downloadable computer software using artificial intelligence for making podcasts and other multimedia content; platform as a service (PaaS) featuring computer software platforms using artificial intelligence for making podcasts and other multimedia content; providing on-line non-downloadable computer software using artificial intelligence for recording, transcribing, editing, and mixing audio, video, text, and other media content; platform as a service (PaaS) featuring computer software platforms using artificial intelligence (AI) for recording, transcribing, editing, and mixing audio, video, text, and other media content; technical support services relating to recording, transcribing, editing, and mixing audio, video, text, and other media content, namely, troubleshooting in the nature of diagnosing computer software problems using artificial intelligence; providing on-line non-downloadable computer software using artificial intelligence for enabling editing of sound files and lyrics in text form; platform as a service (PaaS) featuring computer software and audio and word processing platforms using artificial intelligence for enabling editing of sound files and lyrics in text form; providing on-line non-downloadable computer software using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; platform as a service (PaaS) featuring computer software and audio and word processing platforms using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; providing on-line non-downloadable computer software using artificial intelligence for recording audio and transforming that audio into text, audio, images, video, or combinations thereof; providing on-line non-downloadable computer software using artificial intelligence (AI) for creating avatars; platform as a service (PaaS) featuring computer software and audio and word processing platforms using artificial intelligence for creating avatars.
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
(1) Downloadable and recorded computer software platforms using artificial intelligence for making podcasts and other multimedia content; downloadable and recorded computer software platforms using artificial intelligence for recording, transcribing, editing, and mixing audio, video, text, and other media content; downloadable and recorded audio word processing computer software platforms using artificial intelligence for enabling editing of sound files and lyrics in text form; downloadable computer software using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; downloadable computer software using artificial intelligence for recording audio and transforming that audio into text, audio, images, video, or combinations thereof; downloadable digital image files of avatars; downloadable computer software using artificial intelligence for creating avatars. (1) Providing on-line non-downloadable computer software using artificial intelligence for making podcasts and other multimedia content; platform as a service (PaaS) featuring computer software platforms using artificial intelligence for making podcasts and other multimedia content; providing on-line non-downloadable computer software using artificial intelligence for recording, transcribing, editing, and mixing audio, video, text, and other media content; platform as a service (PaaS) featuring computer software platforms using artificial intelligence (AI) for recording, transcribing, editing, and mixing audio, video, text, and other media content; technical support services relating to recording, transcribing, editing, and mixing audio, video, text, and other media content, namely, troubleshooting in the nature of diagnosing computer software problems using artificial intelligence; providing on-line non-downloadable computer software using artificial intelligence for enabling editing of sound files and lyrics in text form; platform as a service (PaaS) featuring computer software and audio and word processing platforms using artificial intelligence for enabling editing of sound files and lyrics in text form; providing on-line non-downloadable computer software using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; platform as a service (PaaS) featuring computer software and audio and word processing platforms using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; providing on-line non-downloadable computer software using artificial intelligence for recording audio and transforming that audio into text, audio, images, video, or combinations thereof; providing on-line non-downloadable computer software using artificial intelligence (AI) for creating avatars; platform as a service (PaaS) featuring computer software and audio and word processing platforms using artificial intelligence for creating avatars.
6.
APPROACHES TO TRAINING AND IMPLEMENTING A UNIVERSAL VARIABLE MODEL FOR DYNAMIC VOICE SYNTHESIS AND SYSTEMS FOR ACCOMPLISHING THE SAME
Introduced here are approaches to training and then employing computer-implemented models designed to generate synthesized speech using a Universal Variable Model (UVM). The UVM is pre-trained using reference audio samples and associated text prompts to comprehend and replicate various aspects of human speech, including intonation, rhythm, and pronunciation. In the training process, the UVM learns general patterns and relationships between the acoustic properties of speech and the linguistic features of text from a dataset covering different linguistic contexts, accents, and speakers. This enables the UVM to generate natural-sounding speech without the need for personalized training on the user's voice. Users of the media production platform can submit text inputs along with a reference audio sample, and the UVM will produce corresponding audio output in the same voice as the reference sample.
G10L 13/02 - Methods for producing synthetic speechSpeech synthesisers
G10L 13/10 - Prosody rules derived from textStress or intonation
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
7.
APPROACHES TO EDITING AUDIO CONTENT USING DYNAMIC VOICE SYNTHESIS AND SYSTEMS FOR ACCOMPLISHING THE SAME
Introduced here are approaches to editing audio content using dynamic voice synthesis and systems for accomplishing the same. The system uses a transcript associated with an audio file and received input that indicates a location to add or remove text to identify preceding and succeeding segments around the indicated location. The system constructs a modified transcript, and applies a model (e.g., a Universal Variable Model (UVM)) to generate new audio content in accordance with the modified transcript. The model aligns the acoustic properties of the original audio file with linguistic features of the transcript, therefore enabling the new audio content to emulate the original audio file's properties. The system produces a final audio file by inserting the new audio content into the original audio file. This approach allows for dynamic editing of audio content, maintaining coherence and acoustic consistency while accommodating textual modifications. Prior to generating the final audio file, the system can perform one or more authentication operations using a dynamically generated consent statement.
Introduced here are computer programs and associated computer-implemented techniques for manipulating noisy audio signals to produce clean audio signals that are sufficiently high quality so as to be largely, if not entirely, indistinguishable from “rich” recordings generated by recording studios. When a noisy audio signal is obtained by a media production platform, the noisy audio signal can be manipulated to sound as if recording occurred with sophisticated equipment in a soundproof environment. Manipulation can be performed by a model that, when applied to the noisy audio signal, can manipulate its characteristics so as to emulate the characteristics of clean audio signals that are learned through training.
G10L 15/06 - Creation of reference templatesTraining of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
G10L 25/18 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
G10L 25/21 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being power information
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
42 - Scientific, technological and industrial services, research and design
Goods & Services
Research and development in the field of artificial intelligence, and machine learning; Research and development in the field of media, namely, development of software for recording, transcribing, editing, and mixing audio, video, text, and other media content
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable software for creating, recording, editing, mixing, processing, enhancing, and repairing video, audio, text, and other media content; Downloadable computer software using artificial intelligence (AI) for for creating, recording, editing, mixing, processing, enhancing, and repairing video, audio, text, and other media content; Downloadable software for editing audio and dialogue to improve tone; Downloadable computer software using artificial intelligence (AI) for editing audio and dialogue to improve tone; Downloadable software for reducing background noise in video, audio, and other media content; Downloadable computer software using artificial intelligence (AI) for reducing background noise in video, audio, and other media content; Downloadable software for facilitating production of multimedia compilations; Downloadable computer software using artificial intelligence (AI) for facilitating production of multimedia compilations Providing on-line non-downloadable software for creating, recording, editing, mixing, processing, enhancing, and transcribing video and audio; Providing on-line non-downloadable software using artificial intelligence (AI) for creating, recording, editing, mixing, processing, enhancing, and transcribing video and audio; Providing on-line non-downloadable software for for facilitating production of multimedia compilations; Providing on-line non-downloadable software using artificial intelligence (AI) for facilitating production of multimedia compilations; Providing on-line non-downloadable software for reducing background noise in videos and audio; Providing on-line non-downloadable software using artificial intelligence (AI) for reducing background noise in videos and audio
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable software for creating, recording, editing, mixing, processing, enhancing, and transcribing video and audio; Downloadable computer software using artificial intelligence (AI) for creating, recording, editing, mixing, processing, enhancing, and transcribing video and audio; Downloadable software for reducing background noise in videos and audio; Downloadable computer software using artificial intelligence (AI) for reducing background noise in videos and audio; Downloadable software for facilitating production of multimedia compilations; Downloadable computer software using artificial intelligence (AI) for facilitating production of multimedia compilations Providing on-line non-downloadable software for creating, recording, editing, mixing, processing, enhancing, and transcribing video and audio; Providing on-line non-downloadable software using artificial intelligence (AI) for creating, recording, editing, mixing, processing, enhancing, and transcribing video and audio; Providing on-line non-downloadable software for reducing background noise in video, audio, and other media content; Providing on-line non-downloadable software using artificial intelligence (AI) for reducing background noise in video, audio, and other media content; Providing on-line non-downloadable software for for facilitating production of multimedia compilations; Providing on-line non-downloadable software using artificial intelligence (AI) for facilitating production of multimedia compilations
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable software for facilitating production of multimedia compilations; Downloadable computer software using artificial intelligence (AI) for facilitating production of multimedia compilations; Downloadable software for creating, recording, editing, mixing, and processing video; Downloadable computer software using artificial intelligence (AI) for creating, recording, editing, mixing, and processing video Providing on-line non-downloadable software for creating, recording, editing, mixing, and processing video; Providing on-line non-downloadable software for facilitating production of multimedia compilations; Providing on-line non-downloadable software using artificial intelligence (AI) for facilitating production of multimedia compilations; Providing on-line non-downloadable software using artificial intelligence (AI) for creating, recording, editing, mixing, and processing video
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable and recorded computer software platforms using artificial intelligence for making podcasts and other multimedia content; downloadable and recorded computer software platforms using artificial intelligence for recording, transcribing, editing, and mixing audio, video, text, and other media content; downloadable and recorded audio word processing computer software platforms using artificial intelligence for enabling editing of sound files and lyrics in text form; downloadable computer software using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; downloadable computer software using artificial intelligence for recording audio and transforming that audio into text, audio, images, video, or combinations thereof; Downloadable digital image files of avatars; downloadable computer software using artificial intelligence for creating avatars Providing on-line non-downloadable computer software using artificial intelligence for making podcasts and other multimedia content; platform as a service (PAAS) featuring computer software platforms using artificial intelligence for making podcasts and other multimedia content; providing on-line non-downloadable computer software using artificial intelligence for recording, transcribing, editing, and mixing audio, video, text, and other media content; platform as a service (PAAS) featuring computer software platforms using artificial intelligence (AI) for recording, transcribing, editing, and mixing audio, video, text, and other media content; technical support services relating to recording, transcribing, editing, and mixing audio, video, text, and other media content, namely, troubleshooting in the nature of diagnosing computer software problems using artificial intelligence; providing on-line non-downloadable computer software using artificial intelligence for enabling editing of sound files and lyrics in text form; platform as a service (PAAS) featuring computer software and audio and word processing platforms using artificial intelligence for enabling editing of sound files and lyrics in text form; providing on-line non-downloadable computer software using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; platform as a service (PAAS) featuring computer software and audio and word processing platforms using artificial intelligence for use in prompting questions, offering observations, and providing challenges to users in the field of audio, video, text, and other media content; providing on-line non-downloadable computer software using artificial intelligence for recording audio and transforming that audio into text, audio, images, video, or combinations thereof; providing on-line non-downloadable computer software using artificial intelligence (AI) for creating avatars; platform as a service (PAAS) featuring computer software and audio and word processing platforms using artificial intelligence for creating avatars
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable and recorded computer software platforms using artificial intelligence for making podcasts and other audio and video content; downloadable and recorded computer software platforms using artificial intelligence for recording, transcribing, and editing audio, video, and text and for mixing audio and video; downloadable and recorded audio word processing computer software platforms using artificial intelligence for editing of sound files and lyrics in text form; downloadable computer software using artificial intelligence for use in providing analysis and recommendations for improving and editing audio, video, and text; downloadable computer software using artificial intelligence for recording audio and using that audio to create text, audio, images, video, and combinations thereof; downloadable digital image files of avatars; downloadable computer software using artificial intelligence for creating avatars Providing on-line non-downloadable computer software using artificial intelligence for making podcasts and other audio and video content; platform as a service (PAAS) featuring computer software platforms using artificial intelligence for making podcasts and other audio and video content; providing on-line non-downloadable computer software using artificial intelligence for recording, transcribing, and editing audio, video, and text and for mixing audio and video; platform as a service (PAAS) featuring computer software platforms using artificial intelligence (AI) for recording, transcribing, and editing audio, video, and text and for mixing audio and video; technical support services relating to recording, transcribing, editing, and mixing audio, video, text, and other media content, namely, troubleshooting in the nature of diagnosing computer software problems using artificial intelligence; providing on-line non-downloadable computer software using artificial intelligence for editing of sound files and lyrics in text form; platform as a service (PAAS) featuring computer software and audio and word processing platforms using artificial intelligence for enabling editing of sound files and lyrics in text form; providing on-line non-downloadable computer software using artificial intelligence for use in providing analysis and recommendations for improving and editing audio, video, and text; platform as a service (PAAS) featuring computer software and audio and word processing platforms using artificial intelligence for use in providing analysis and recommendations for improving and editing audio, video, and text; providing online non-downloadable computer software using artificial intelligence for recording audio and using that audio to create text, audio, images, video, and combinations thereof; providing on-line non-downloadable computer software using artificial intelligence (AI) for creating avatars; platform as a service (PAAS) featuring computer software and audio and word processing platforms using artificial intelligence for creating avatars
15.
SIMULTANEOUS RECORDING AND UPLOADING OF MULTIPLE AUDIO FILES OF THE SAME CONVERSATION AND AUDIO DRIFT NORMALIZATION SYSTEMS AND METHODS
As advances in Internet communications including video and audio mediums continue to increase, the need to effectively and efficiently communicate media has also increased. Many of the platforms that are designed for the creation and communication of media content cannot ensure that the media content is accessible in the event of a problem, for example, with the recording process. Moreover, these platforms generally do not eliminate the need to record a complete session in its entirety prior to uploading or using that content. Introduced here is an approach to content recording that involves two computer programs. A first computer program executing on a first device can segment media into chunks as the media is being recorded and then transmit those chunks, one by one, to a second device. A second computer program executing on the second device can then reconstruct the chunks in numbered sequence.
Different types of media experiences can be developed based on characteristics of the consumer. “Linear” experiences may require execution of a pre-built script, although the script could be dynamically modified by a media production platform. Linear experiences can include guided audio tours that are modified or updated based on the location of the consumer. “Enhanced” experiences include conventional media content that is supplemented with intelligent media content. For example, turn-by-turn directions could be supplemented with audio descriptions about the surrounding area. “Freeform” experiences, meanwhile, are those that can continually morph based on information gleaned from a consumer. For example, a radio station may modify what content is being presented based on the geographical metadata uploaded by a computing device associated with the consumer.
G06F 3/0484 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range
G06F 3/04817 - Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance using icons
G06F 16/68 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
G06F 16/683 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
G10L 15/187 - Phonemic context, e.g. pronunciation rules, phonotactical constraints or phoneme n-grams
Introduced here are approaches to training and then employing computer-implemented models designed to upsample discrete audio signals to higher sampling rates. Assume, for example, that a media production platform obtains a first discrete signal at a relatively low sampling rate. The relatively low sampling frequency may make the first discrete audio signal unsuitable for inclusion in media compilations, so the media production platform may attempt to improve its quality through upsampling. To accomplish this, the media production platform can apply a transform to the first discrete signal to produce a first magnitude spectrogram. Then, the media production platform can apply a computer-implemented model to the first magnitude spectrogram to produce a second magnitude spectrogram. Thereafter, the media production platform can apply an inverse transform to the second magnitude spectrogram to create a second discrete signal that has a higher sampling rate than the first discrete audio signal.
G10L 25/18 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
G06N 3/088 - Non-supervised learning, e.g. competitive learning
G10L 19/02 - Speech or audio signal analysis-synthesis techniques for redundancy reduction, e.g. in vocodersCoding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
18.
TRAINING GENERATIVE ADVERSARIAL NETWORKS TO UPSAMPLE AUDIO
Introduced here are approaches to training and then employing computer-implemented models designed to upsample discrete audio signals to higher sampling rates. Assume, for example, that a media production platform obtains a first discrete signal at a relatively low sampling rate. The relatively low sampling frequency may make the first discrete audio signal unsuitable for inclusion in media compilations, so the media production platform may attempt to improve its quality through upsampling. To accomplish this, the media production platform can apply a transform to the first discrete signal to produce a first magnitude spectrogram. Then, the media production platform can apply a computer-implemented model to the first magnitude spectrogram to produce a second magnitude spectrogram. Thereafter, the media production platform can apply an inverse transform to the second magnitude spectrogram to create a second discrete signal that has a higher sampling rate than the first discrete audio signal.
G10L 25/18 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
G06N 3/088 - Non-supervised learning, e.g. competitive learning
G10L 19/02 - Speech or audio signal analysis-synthesis techniques for redundancy reduction, e.g. in vocodersCoding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
19.
FILLER WORD DETECTION THROUGH TOKENIZING AND LABELING OF TRANSCRIPTS
Introduced here are computer programs and associated computer-implemented techniques for discovering the presence of filler words through tokenization of a transcript derived from audio content. When audio content is obtained by a media production platform, the audio content can be converted into text content as part of a speech-to-text operation. The text content can then be tokenized and labeled using a Natural Language Processing (NLP) library. Tokenizing/labeling may be performed in accordance with a series of rules associated with filler words. At a high level, these rules may examine the text content (and associated tokens/labels) to determine whether patterns, relationships, verbatim, and context indicate that a term is a filler word. Any filler words that are discovered in the text content can be identified as such so that appropriate action(s) can be taken.
Introduced here are computer programs and associated computer-implemented techniques for facilitating the creation of a master transcription (or simply “transcript”) that more accurately reflects underlying audio by comparing multiple independently generated transcripts. The master transcript may be used to record and/or produce various forms of media content, as further discussed below. Thus, the technology described herein may be used to facilitate editing of text content, audio content, or video content. These computer programs may be supported by a media production platform that is able to generate the interfaces through which individuals (also referred to as “users”) can create, edit, or view media content. For example, a computer program may be embodied as a word processor that allows individuals to edit voice-based audio content by editing a master transcript, and vice versa.
Media content can be created and/or modified using a network-accessible platform. Scripts for content-based experiences could be readily created using one or more interfaces generated by the network-accessible platform. For example, a script for a content-based experience could be created using an interface that permits triggers to be inserted directly into the script. Interface(s) may also allow different media formats to be easily aligned for post-processing. For example, a transcript and an audio file may be dynamically aligned so that the network-accessible platform can globally reflect changes made to either item. User feedback may also be presented directly on the interface(s) so that modifications can be made based on actual user experiences.
G06Q 10/101 - Collaborative creation, e.g. joint development of products or services
G10L 19/008 - Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
G11B 27/031 - Electronic editing of digitised analogue information signals, e.g. audio or video signals
G11B 27/10 - IndexingAddressingTiming or synchronisingMeasuring tape travel
Media content can be created and/or modified using a network-accessible platform. Scripts for content-based experiences could be readily created using one or more interfaces generated by the network-accessible platform. For example, a script for a content-based experience could be created using an interface that permits triggers to be inserted directly into the script. Interface(s) may also allow different media formats to be easily aligned for post-processing. For example, a transcript and an audio file may be dynamically aligned so that the network-accessible platform can globally reflect changes made to either item. User feedback may also be presented directly on the interface(s) so that modifications can be made based on actual user experiences.
G06Q 10/101 - Collaborative creation, e.g. joint development of products or services
G10L 19/008 - Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
G11B 27/031 - Electronic editing of digitised analogue information signals, e.g. audio or video signals
G11B 27/10 - IndexingAddressingTiming or synchronisingMeasuring tape travel
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable computer software that utilizes artificial
intelligence (AI) to assist in creating content through a
software platform for recording, transcribing, editing, and
mixing audio, video, text and other media content, the AI
assisting in content creation; downloadable computer
software that utilizes artificial intelligence to engage
users of a software platform through prompting questions,
offering observations, and providing challenges, the
software platform allowing the users to record, transcribe,
edit, and mix audio, video, text, and other media content;
downloadable computer software that utilizes artificial
intelligence to simulate stimulus provided by a
conversationalist, guiding a writer's creative flow and
nurturing the writer's ability to flesh out and structure
ideas effectively; downloadable files of avatars for use in
virtual environments and for use in software platforms for
recording, transcribing, editing, and mixing audio, video,
text, and other media content; downloadable computer
software for making audio, video and text content, in a
semi- or fully-autonomous manner; downloadable computer
software for facilitating the recording, transcribing,
editing, and mixing of audio, video, text, and other media
content; downloadable computer software for transforming
ideas specified, either audibly or textually, by an
individual into usable outputs in an automated manner, while
also allowing the individual to edit those outputs for the
purpose of producing content; downloadable word processor
computer programs for enabling individuals to edit audio
through the manipulation of corresponding text, and vice
versa; downloadable computer programs that enable
individuals to edit audio and lyrics in text form;
downloadable computer software for artificial intelligence;
downloadable computer software for avatars; downloadable
computer software through which an individual is able to
audibly or textually record thoughts and receive text,
audio, images, video, or combinations thereof that are
produced as output; downloadable computer software through
which inputs are identified, provided, or generated in
audible, visual, or textual form and those inputs are used
to guide identification or generation of outputs in audible,
visual, or textual form. Providing non-downloadable computer software that utilizes
artificial intelligence (AI) to assist in creating content
through a software platform for recording, transcribing,
editing, and mixing audio, video, text and other media
content, the AI assisting in content creation; providing
non-downloadable computer software that utilizes artificial
intelligence to engage users of a software platform through
prompting questions, offering observations, and providing
challenges, the software platform allowing the users to
record, transcribe, edit, and mix audio, video, text, and
other media content; providing non-downloadable computer
software that utilizes artificial intelligence to simulate
stimulus provided by a conversationalist, guiding a writer's
creative flow and nurturing the writer's ability to flesh
out and structure ideas effectively; providing
non-downloadable computer software for creating avatars for
use in virtual environments and for use in software
platforms for recording, transcribing, editing, and mixing
audio, video, text, and other media content; providing
non-downloadable computer software for making audio, video
and text content, in a semi- or fully-autonomous manner;
providing non-downloadable computer software for
facilitating the recording, transcribing, editing, and
mixing of audio, video, text, and other media content;
providing non-downloadable computer software for
transforming ideas specified, either audibly or textually,
by an individual into usable outputs in an automated manner,
while also allowing the individual to edit those outputs for
the purpose of producing content; providing non-downloadable
word processor computer programs for enabling individuals to
edit audio through the manipulation of corresponding text,
and vice versa; providing non-downloadable computer programs
that enable individuals to edit audio and lyrics in text
form; providing non-downloadable computer software for
artificial intelligence; providing non-downloadable computer
software for avatars; providing non-downloadable computer
software through which an individual is able to audibly or
textually record thoughts and receive text, audio, images,
video, or combinations thereof that are produced as output;
providing non-downloadable computer software through which
inputs are identified, provided, or generated in audible,
visual, or textual form and those inputs are used to guide
identification or generation of outputs in audible, visual,
or textual form.
24.
AUTOMATED GENERATION OF TRANSCRIPTS THROUGH INDEPENDENT TRANSCRIPTION
Introduced here are computer programs and associated computer-implemented techniques for facilitating the creation of a master transcription (or simply “transcript”) that more accurately reflects underlying audio by comparing multiple independently generated transcripts. The master transcript may be used to record and/or produce various forms of media content, as further discussed below. Thus, the technology described herein may be used to facilitate editing of text content, audio content, or video content. These computer programs may be supported by a media production platform that is able to generate the interfaces through which individuals (also referred to as “users”) can create, edit, or view media content. For example, a computer program may be embodied as a word processor that allows individuals to edit voice-based audio content by editing a master transcript, and vice versa.
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable computer software for providing artificial
intelligence (AI) integrated into software platforms for
recording, transcribing, editing, and mixing audio, video,
text and other media content, the AI assisting in content
creation; downloadable computer software for providing
artificial intelligence (AI) integrated into software
platforms for recording, transcribing, editing, and mixing
audio, video, text and other media content, the AI engaging
users through prompting questions, offering observations,
and providing constructive challenges; downloadable software
providing artificial intelligence which simulates stimulus
provided by a great conversationalist, guiding a writer's
creative flow and nurturing the writer's ability to flesh
out and structure their ideas effectively; downloadable
image files of avatars for use in virtual environments and
for use in platforms for recording, transcribing, editing,
and mixing audio, video, text and other media content;
downloadable computer software platforms for making audio,
video and text content; downloadable computer software
platforms for recording, transcribing, editing, and mixing
audio, video, text and other media content; downloadable
audio word processing computer software platforms enabling
editors and producers to edit sound files and writers to
edit lyrics in text form; downloadable computer software for
artificial intelligence; downloadable computer software for
avatars. Application service provider for providing artificial
intelligence (AI) integrated into platform as a service for
recording, transcribing, editing, and mixing audio, video,
text and other media content, the AI assisting in content
creation; application service provider for providing
artificial intelligence (AI) integrated into platform as a
service for recording, transcribing, editing, and mixing
audio, video, text and other media content, the AI engaging
users through prompting questions, offering observations,
and providing constructive challenges; application service
provider providing artificial intelligence (AI) integrated
into platform as a service providing artificial intelligence
which simulates stimulus provided by a great
conversationalist, guiding a writer's creative flow and
nurturing the writer's ability to flesh out and structure
their ideas effectively; software as a service providing
avatars for use in virtual environments and for use in
platforms for recording, transcribing, editing, and mixing
audio, video, text and other media content; software as a
service for providing artificial intelligence; software as a
service for providing avatars.
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
(1) Downloadable computer software that utilizes artificial intelligence (AI) to assist in creating content through a software platform for recording, transcribing, editing, and mixing audio, video, text, images and graphics, the AI assisting in content creation; downloadable computer software that utilizes artificial intelligence to engage users of a software platform through prompting questions, offering observations, and providing challenges, the software platform allowing the users to record, transcribe, edit, and mix audio, video, text, images and graphics; downloadable computer software that utilizes artificial intelligence to simulate stimulus provided by a conversationalist, guiding a writer's creative flow and nurturing the writer's ability to flesh out and structure ideas effectively; downloadable files of avatars for use in virtual environments and for use in software platforms for recording, transcribing, editing, and mixing audio, video, text, images and graphics; downloadable computer software for making audio, video and text content, in a semi- or fully-autonomous manner; downloadable computer software for facilitating the recording, transcribing, editing, and mixing of audio, video, text, images and graphics; downloadable computer software for transforming ideas specified, either audibly or textually, by an individual into usable outputs in an automated manner, while also allowing the individual to edit those outputs for the purpose of producing content; downloadable word processor computer programs for enabling individuals to edit audio through the manipulation of corresponding text, and vice versa; downloadable computer programs that enable individuals to edit audio and lyrics in text form; downloadable computer software that utilizes artificial intelligence for providing feedback in the field of writing; downloadable computer software for creating avatars in virtual worlds; downloadable computer software through which an individual is able to audibly or textually record thoughts and receive text, audio, images, video, or combinations thereof that are produced as output; downloadable computer software through which inputs are identified, provided, or generated in audible, visual, or textual form and those inputs are used to guide identification or generation of outputs in audible, visual, or textual form. (1) Providing non-downloadable computer software that utilizes artificial intelligence (AI) to assist in creating content through a software platform for recording, transcribing, editing, and mixing audio, video, text, images and graphics, the AI assisting in content creation; providing non-downloadable computer software that utilizes artificial intelligence to engage users of a software platform through prompting questions, offering observations, and providing challenges, the software platform allowing the users to record, transcribe, edit, and mix audio, video, text, images and graphics; providing non-downloadable computer software that utilizes artificial intelligence to simulate stimulus provided by a conversationalist, guiding a writer's creative flow and nurturing the writer's ability to flesh out and structure ideas effectively; providing non-downloadable computer software for creating avatars for use in virtual environments and for use in software platforms for recording, transcribing, editing, and mixing audio, video, text, images and graphics; providing non-downloadable computer software for making audio, video and text content, in a semi- or fully-autonomous manner; providing non-downloadable computer software for facilitating the recording, transcribing, editing, and mixing of audio, video, text, images and graphics; providing non-downloadable computer software for transforming ideas specified, either audibly or textually, by an individual into usable outputs in an automated manner, while also allowing the individual to edit those outputs for the purpose of producing content; providing non-downloadable word processor computer programs for enabling individuals to edit audio through the manipulation of corresponding text, and vice versa; providing non-downloadable computer programs that enable individuals to edit audio and lyrics in text form; providing non-downloadable computer software that utilizes artificial intelligence for providing feedback in the field of writing; providing non-downloadable computer software for creating avatars in virtual worlds; providing non-downloadable computer software through which an individual is able to audibly or textually record thoughts and receive text, audio, images, video, or combinations thereof that are produced as output; providing non-downloadable computer software through which inputs are identified, provided, or generated in audible, visual, or textual form and those inputs are used to guide identification or generation of outputs in audible, visual, or textual form.
09 - Scientific and electric apparatus and instruments
Goods & Services
Downloadable computer software that utilizes artificial intelligence (AI) for use in creating content, namely, for recording, transcribing, editing, and mixing audio, video, text and other media content; downloadable computer software that utilizes artificial intelligence (AI) for use in prompting questions, offering observations, and providing challenges to users in the field of recording, transcribing, editing, and mixing audio, video, text, and other media content; downloadable computer software that utilizes artificial intelligence (AI) for use in providing feedback in the field of writing; downloadable image files of avatars for use in virtual worlds; downloadable computer software using artificial intelligence (AI) for use in creating and editing media content and producing media compilations containing audio, video and text content; downloadable computer software for recording, transcribing, editing, and mixing audio, video, text, and other media content; downloadable computer software using artificial intelligence (AI) for use in transforming and editing audio and text into multimedia compilations; downloadable computer programs for word processing, namely, for editing corresponding text and audio; downloadable computer programs for editing audio and lyrics in text form; downloadable computer software using artificial intelligence (AI) for creating and editing media content and producing media compilations; downloadable computer software for creating avatars in virtual worlds; downloadable computer software for recording audio and transforming that audio into text, audio, images, video, or combinations thereof; downloadable computer software for recording, processing, and editing audio, images, video, and text
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable computer software that utilizes artificial intelligence (AI) to assist in creating content through a software platform for recording, transcribing, editing, and mixing audio, video, text and other media content, the AI assisting in content creation in the field of media editing; downloadable computer software that utilizes artificial intelligence to engage users of a software platform through prompting questions, offering observations, and providing challenges, the software platform allowing the users to record, transcribe, edit, and mix audio, video, text, and other media content in the field of media editing; downloadable computer software that utilizes artificial intelligence to simulate stimulus provided by a conversationalist, guiding a writer's creative flow and nurturing the writer's ability to flesh out and structure ideas effectively in the field of media editing; downloadable digital image files of avatars for use in virtual environments and for use in software platforms in connection with downloadable software for recording, transcribing, editing, and mixing audio, video, text, and other media content in the field of media editing; downloadable computer software using artificial intelligence for creating, editing, or compiling audio, video and text content, in a semi- or fully-autonomous manner; downloadable computer software for facilitating the recording, transcribing, editing, and mixing of audio, video, text, and other media content in the field of media editing; downloadable computer software for transforming ideas specified, either audibly or textually, by an individual into usable outputs in an automated manner, while also allowing the individual to edit those outputs for the purpose of producing content, namely, for use in facilitating production of multimedia compilations; downloadable word processor computer programs for enabling individuals to edit audio through the manipulation of corresponding text, and vice versa; downloadable computer programs that enable individuals to edit audio and lyrics in text form; downloadable computer software that uses artificial intelligence for facilitating production of multimedia compilations; downloadable computer software for creating avatars in virtual worlds; downloadable computer software using artificial intelligence through which an individual is able to audibly or textually record thoughts and receive text, audio, images, video, or combinations thereof that are produced as output for creating multimedia compilations; downloadable computer software using artificial intelligence through which inputs are identified, provided, or generated in audible, visual, or textual form and those inputs are used to guide identification or generation of outputs in audible, visual, or textual form for facilitating production of multimedia compilations Providing temporary use of online non-downloadable computer software that utilizes artificial intelligence (AI) to assist in creating content through a software platform for recording, transcribing, editing, and mixing audio, video, text and other media content, the AI assisting in content creation in the field of media editing; providing temporary use of online non-downloadable computer software that utilizes artificial intelligence to engage users of a software platform through prompting questions, offering observations, and providing challenges, the software platform allowing the users to record, transcribe, edit, and mix audio, video, text, and other media content in the field of media editing; providing temporary use of online non-downloadable computer software that utilizes artificial intelligence to simulate stimulus provided by a conversationalist, guiding a writer's creative flow and nurturing the writer's ability to flesh out and structure ideas effectively in the field of media editing; providing temporary use of online non-downloadable files of avatars for use in virtual environments and for use in software platforms in connection with downloadable software for recording, transcribing, editing, and mixing audio, video, text, and other media content in the field of media editing; providing temporary use of online non-downloadable computer software for creating, editing, or compiling audio, video and text content, in a semi- or fully-autonomous manner; providing temporary use of online non-downloadable computer software for facilitating the recording, transcribing, editing, and mixing of audio, video, text, and other media content in the field of media editing; providing temporary use of online non-downloadable computer software for transforming ideas specified, either audibly or textually, by an individual into usable outputs in an automated manner, while also allowing the individual to edit those outputs for the purpose of producing content, namely, for use in facilitating production of multimedia compilations; providing temporary use of online non-downloadable word processor computer programs for enabling individuals to edit audio through the manipulation of corresponding text, and vice versa; providing temporary use of online non-downloadable computer programs that enable individuals to edit audio and lyrics in text form; providing temporary use of online non-downloadable computer software for creating avatars in virtual worlds; providing temporary use of online non-downloadable computer software using artificial intelligence through which an individual is able to audibly or textually record thoughts and receive text, audio, images, video, or combinations thereof that are produced as output for creating multimedia compilations; providing temporary use of online non-downloadable computer software using artificial intelligence through which inputs are identified, provided, or generated in audible, visual, or textual form and those inputs are used to guide identification or generation of outputs in audible, visual, or textual form for facilitating production of multimedia compilations
42 - Scientific, technological and industrial services, research and design
Goods & Services
Providing temporary use of online non-downloadable computer software that utilizes artificial intelligence (AI) for use in creating content, namely, for recording, transcribing, editing, and mixing audio, video, text and other media content; providing temporary use of online non-downloadable computer software that utilizes artificial intelligence (AI) for use in prompting questions, offering observations, and providing challenges to users in the field of recording, transcribing, editing, and mixing audio, video, text, and other media content; providing temporary use of online non-downloadable computer software that utilizes artificial intelligence (AI) for use in providing feedback in the field of writing; providing temporary use of online non-downloadable image files of avatars for use in virtual worlds; providing temporary use of online non-downloadable computer software using artificial intelligence (AI) for use in creating and editing media content and producing media compilations containing audio, video and text content; providing temporary use of online non-downloadable computer software for recording, transcribing, editing, and mixing audio, video, text, and other media content; providing temporary use of online non-downloadable computer software using artificial intelligence (AI) for use in transforming and editing audio and text into multimedia compilations; providing temporary use of online non-downloadable computer programs for word processing, namely, for editing corresponding text and audio; providing temporary use of online non-downloadable computer programs for editing audio and lyrics in text form; providing temporary use of online non-downloadable computer software using artificial intelligence (AI) for creating and editing media content and producing media compilations; providing temporary use of online non-downloadable computer software for creating avatars in virtual worlds; providing temporary use of online non-downloadable computer software for recording audio and transforming that audio into text, audio, images, video, or combinations thereof; providing temporary use of online non-downloadable computer software for recording, processing, and editing audio, images, video, and text
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
(1) Downloadable computer software that utilizes artificial intelligence (AI) integrated into software platforms for recording, transcribing, editing, and mixing audio, video, text, images, graphics, the AI assisting in creating audio-visual media content; downloadable computer software that utilizes artificial intelligence (AI) integrated into software platforms in the field of audio and video editing allowing the users to record, transcribe, edit, and mix audio, video, text, with the AI engaging users through prompting questions, offering observations, and providing constructive challenges in the field of audio and video editing; downloadable software providing artificial intelligence which simulates stimulus provided by a great conversationalist, guiding a writer's creative flow and nurturing the writer's ability to flesh out and structure their ideas effectively; downloadable image files of avatars for use in virtual environments and for use in platforms for recording, transcribing, editing, and mixing audio, video, text and other media content; downloadable computer software platforms for recording, transcribing, editing, and mixing of audio, video, and text for virtual computer games; downloadable computer software platforms for recording, transcribing, editing, and mixing audio, video, text, images and graphics; downloadable audio word processing computer software platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form; downloadable computer software using artificial intelligence for audio and video editing; downloadable computer software for use in creating avatars in virtual worlds. (1) Application service provider (ASP) providing computer software applications of others, such computer software applications utilizing artificial intelligence (AI) integrated into software platforms for recording, transcribing, editing, and mixing audio, video, text, images, graphics, the AI assisting in creating audio-visual media content; Application service provider (ASP), providing computer software applications of others featuring computer software that utilizes artificial intelligence (AI) integrated into software platforms in the field of audio and video editing, allowing the users to record, transcribe, edit, and mix audio, video, and text, with the AI engaging users through prompting questions, offering observations, and providing constructive challenges in the field of audio and video editing; application service provider providing artificial intelligence (AI) integrated into platform as a service providing artificial intelligence which simulates stimulus provided by a great conversationalist, guiding a writer's creative flow and nurturing the writer's ability to flesh out and structure their ideas effectively; software as a service providing avatars for use in virtual environments and for use in platforms for recording, transcribing, editing, and mixing audio, video, text and other media content; Software as a service (SAAS) services offering computer software using artificial intelligence for audio and video editing; Software as a service (SAAS) services offering software for use in creating avatars in virtual worlds.
31.
Tokenization of text data to facilitate automated discovery of speech disfluencies
Introduced here are computer programs and associated computer-implemented techniques for discovering the presence of filler words through tokenization of a transcript derived from audio content. When audio content is obtained by a media production platform, the audio content can be converted into text content as part of a speech-to-text operation. The text content can then be tokenized and labeled using a Natural Language Processing (NLP) library. Tokenizing/labeling may be performed in accordance with a series of rules associated with filler words. At a high level, these rules may examine the text content (and associated tokens/labels) to determine whether patterns, relationships, verbatim, and context indicate that a term is a filler word. Any filler words that are discovered in the text content can be identified as such so that appropriate action(s) can be taken.
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable computer software that utilizes artificial intelligence (AI) to assist in creating content, namely, content for social media platforms and other distribution channels through a software platform for recording, transcribing, editing, and mixing audio, video, and text, with the AI assisting in content creation; downloadable computer software that utilizes artificial intelligence (AI) integrated into software platforms in the field of audio and video editing allowing the users to record, transcribe, edit, and mix audio, video, and text, with the AI engaging users through prompting questions, offering observations, and providing constructive challenges in the field of audio and video editing; downloadable computer software that utilizes artificial intelligence to simulate stimulus provided by a conversationalist in the field of writing, guiding a writer's creative flow and nurturing the writer's ability to flesh out and structure ideas effectively; downloadable image files of avatars for use in virtual environments and for use in software platforms for recording, transcribing, editing, and mixing audio, video, text, and other media content for customizing avatars; downloadable computer software platforms for recording, transcribing, editing, and mixing of audio, video, text, and other media content for virtual computer games; downloadable audio word processing computer software platforms for enabling editors and producers to edit sound files and writers to edit lyrics through the manipulation of corresponding text format; downloadable computer software using artificial intelligence for audio and video editing; downloadable computer software for generating avatars for virtual computer environments Application service provider (ASP), namely, hosting computer software applications of others featuring computer software that utilizes artificial intelligence (AI) to assist in creating content, namely, content for social media platforms and other distribution channels through a software platform for recording, transcribing, editing, and mixing audio, video, and text, with the AI assisting in content creation; Application service provider (ASP), namely, hosting computer software applications of others featuring computer software that utilizes artificial intelligence (AI) integrated into software platforms in the field of audio and video editing, allowing the users to record, transcribe, edit, and mix audio, video, and text, with the AI engaging users through prompting questions, offering observations, and providing constructive challenges in the field of audio and video editing; Application service provider (ASP), namely, hosting computer software applications of others featuring non-downloadable computer software that utilizes artificial intelligence to simulate stimulus provided by a conversationalist in the field of writing, guiding a writer's creative flow and nurturing the writer's ability to flesh out and structure ideas effectively; Software as a service (SAAS) services featuring non-downloadable image files of avatars for use in virtual environments and for use in software platforms for recording, transcribing, editing, and mixing audio, video, text, and other media content for customizing avatars; Software as a service (SAAS) services featuring computer software using artificial intelligence for audio and video editing; Software as a service (SAAS) services featuring software for generating avatars for virtual computer environments
33.
Filler word detection through tokenizing and labeling of transcripts
Introduced here are computer programs and associated computer-implemented techniques for discovering the presence of filler words through tokenization of a transcript derived from audio content. When audio content is obtained by a media production platform, the audio content can be converted into text content as part of a speech-to-text operation. The text content can then be tokenized and labeled using a Natural Language Processing (NLP) library. Tokenizing/labeling may be performed in accordance with a series of rules associated with filler words. At a high level, these rules may examine the text content (and associated tokens/labels) to determine whether patterns, relationships, verbatim, and context indicate that a term is a filler word. Any filler words that are discovered in the text content can be identified as such so that appropriate action(s) can be taken.
Introduced here are computer programs and associated computer-implemented techniques for manipulating noisy audio signals to produce clean audio signals that are sufficiently high quality so as to be largely, if not entirely, indistinguishable from “rich” recordings generated by recording studios. When a noisy audio signal is obtained by a media production platform, the noisy audio signal can be manipulated to sound as if recording occurred with sophisticated equipment in a soundproof environment. Manipulation can be performed by a model that, when applied to the noisy audio signal, can manipulate its characteristics so as to emulate the characteristics of clean audio signals that are learned through training.
G10L 15/06 - Creation of reference templatesTraining of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
G10L 25/18 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
G10L 25/21 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being power information
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
35.
APPROACHES TO GENERATING STUDIO-QUALITY RECORDINGS THROUGH MANIPULATION OF NOISY AUDIO
Introduced here are computer programs and associated computer-implemented techniques for manipulating noisy audio signals to produce clean audio signals that are sufficiently high quality so as to be largely, if not entirely, indistinguishable from “rich” recordings generated by recording studios. When a noisy audio signal is obtained by a media production platform, the noisy audio signal can be manipulated to sound as if recording occurred with sophisticated equipment in a soundproof environment. Manipulation can be performed by a model that, when applied to the noisy audio signal, can manipulate its characteristics so as to emulate the characteristics of clean audio signals that are learned through training.
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
G10L 25/18 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
G10L 25/21 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being power information
Different types of media experiences can be developed based on characteristics of the consumer. “Linear” experiences may require execution of a pre-built script, although the script could be dynamically modified by a media production platform. Linear experiences can include guided audio tours that are modified or updated based on the location of the consumer. “Enhanced” experiences include conventional media content that is supplemented with intelligent media content. For example, turn-by-turn directions could be supplemented with audio descriptions about the surrounding area. “Freeform” experiences, meanwhile, are those that can continually morph based on information gleaned from a consumer. For example, a radio station may modify what content is being presented based on the geographical metadata uploaded by a computing device associated with the consumer.
G06F 17/00 - Digital computing or data processing equipment or methods, specially adapted for specific functions
G06F 3/04817 - Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance using icons
G06F 3/0484 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range
G06F 16/68 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
G06F 16/683 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
G10L 15/187 - Phonemic context, e.g. pronunciation rules, phonotactical constraints or phoneme n-grams
The invention relates to simultaneous recording and uploading systems and methods, and, more particularly to a simultaneous recording and uploading multiple files from the same conversation.
Media content can be created and/or modified using a network-accessible platform. Scripts for content-based experiences could be readily created using one or more interfaces generated by the network-accessible platform. For example, a script for a content-based experience could be created using an interface that permits triggers to be inserted directly into the script. Interface(s) may also allow different media formats to be easily aligned for post-processing. For example, a transcript and an audio file may be dynamically aligned so that the network-accessible platform can globally reflect changes made to either item. User feedback may also be presented directly on the interface(s) so that modifications can be made based on actual user experiences.
G06Q 10/101 - Collaborative creation, e.g. joint development of products or services
G10L 19/008 - Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
G11B 27/031 - Electronic editing of digitised analogue information signals, e.g. audio or video signals
G11B 27/10 - IndexingAddressingTiming or synchronisingMeasuring tape travel
The invention relates to audio drift normalization, and more particularly to audio drift normalization systems and methods that can normalize audio drift of a plurality of recordings from a source.
Different types of media experiences can be developed based on characteristics of the consumer. “Linear” experiences may require execution of a pre-built script, although the script could be dynamically modified by a media production platform. Linear experiences can include guided audio tours that are modified or updated based on the location of the consumer. “Enhanced” experiences include conventional media content that is supplemented with intelligent media content. For example, turn-by-turn directions could be supplemented with audio descriptions about the surrounding area. “Freeform” experiences, meanwhile, are those that can continually morph based on information gleaned from a consumer. For example, a radio station may modify what content is being presented based on the geographical metadata uploaded by a computing device associated with the consumer.
G06F 17/00 - Digital computing or data processing equipment or methods, specially adapted for specific functions
G06F 3/0484 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range
G10L 15/187 - Phonemic context, e.g. pronunciation rules, phonotactical constraints or phoneme n-grams
G06F 3/04817 - Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance using icons
G06F 16/68 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
G06F 16/683 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
09 - Scientific and electric apparatus and instruments
35 - Advertising and business services
41 - Education, entertainment, sporting and cultural services
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable computer software platforms for making podcasts and other audio content; downloadable computer software platforms for recording, transcribing, editing, and mixing podcasts and other media content; downloadable audio word processing computer software platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form; downloadable computer software platforms for multimedia production. Providing business support services in the nature of start-up support for businesses of others related to podcasting and podcasting creation; providing business support services in the nature of multimedia production. Audio recording and production services, namely, making podcasts and other audio content; providing a website featuring blogs in the field of podcasting, podcasting creation, and audio production. Providing temporary use of on-line non-downloadable computer software for making podcasts and other audio content; platform as a service (PAAS) featuring computer software platforms for making podcasts and other audio content; providing temporary use of on-line non-downloadable computer software for recording, transcribing, editing, and mixing podcasts and other media content; platform as a service (PAAS) featuring computer software platforms for recording, transcribing, editing, and mixing podcasts and other media content; technical support services related to podcasting and podcasting creation, namely, troubleshooting in the nature of diagnosing computer software problems in the recording, creating, and editing of media; providing temporary use of on-line non-downloadable computer software enabling editors and producers to edit sound files and writers to edit lyrics in text form; platform as a service (PAAS) featuring computer software and audio and word processing platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form; platform as a service (PAAS) featuring computer software platforms for multimedia production.
09 - Scientific and electric apparatus and instruments
35 - Advertising and business services
41 - Education, entertainment, sporting and cultural services
42 - Scientific, technological and industrial services, research and design
Goods & Services
(1) Downloadable computer software platforms for making podcasts and other audio content; downloadable computer software platforms for recording, transcribing, editing, and mixing podcasts and audio-visual media content; downloadable audio word processing computer software platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form; downloadable computer software platforms for podcasts, film, video, television shows, music and radio production. (1) Providing business support services in the nature of start-up support for businesses of others related to podcasting and podcasting creation; providing business support services in the nature of multimedia production.
(2) Audio recording and production services, namely, making podcasts and other audio content; providing blogs in the field of podcasting, podcasting creation, and audio production via a website.
(3) Providing temporary use of on-line non-downloadable computer software for making podcasts and other audio content; platform as a service (PAAS) featuring computer software platforms for making podcasts and other audio content; providing temporary use of on-line non-downloadable computer software for recording, transcribing, editing, and mixing podcasts and audio-visual media content; platform as a service (PAAS) featuring computer software platforms for recording, transcribing, editing, and mixing podcasts and audio-visual media content; technical support services related to podcasting and podcasting creation, namely, troubleshooting in the nature of diagnosing computer software problems in the recording, creating, and editing of media; providing temporary use of on-line non-downloadable computer software enabling editors and producers to edit sound files and writers to edit lyrics in text form; platform as a service (PAAS) featuring computer software and audio and word processing platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form; platform as a service (PAAS) featuring computer software platforms for podcasts, film, video, television shows, music and radio production.
43.
Upsampling of audio using generative adversarial networks
Introduced here are approaches to training and then employing computer-implemented models designed to upsample discrete audio signals to higher sampling rates. Assume, for example, that a media production platform obtains a first discrete signal at a relatively low sampling rate. The relatively low sampling frequency may make the first discrete audio signal unsuitable for inclusion in media compilations, so the media production platform may attempt to improve its quality through upsampling. To accomplish this, the media production platform can apply a transform to the first discrete signal to produce a first magnitude spectrogram. Then, the media production platform can apply a computer-implemented model to the first magnitude spectrogram to produce a second magnitude spectrogram. Thereafter, the media production platform can apply an inverse transform to the second magnitude spectrogram to create a second discrete signal that has a higher sampling rate than the first discrete audio signal.
G10L 25/18 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
G06N 3/088 - Non-supervised learning, e.g. competitive learning
G10L 19/02 - Speech or audio signal analysis-synthesis techniques for redundancy reduction, e.g. in vocodersCoding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
44.
Training generative adversarial networks to upsample audio
Introduced here are approaches to training and then employing computer-implemented models designed to upsample discrete audio signals to higher sampling rates. Assume, for example, that a media production platform obtains a first discrete signal at a relatively low sampling rate. The relatively low sampling frequency may make the first discrete audio signal unsuitable for inclusion in media compilations, so the media production platform may attempt to improve its quality through upsampling. To accomplish this, the media production platform can apply a transform to the first discrete signal to produce a first magnitude spectrogram. Then, the media production platform can apply a computer-implemented model to the first magnitude spectrogram to produce a second magnitude spectrogram. Thereafter, the media production platform can apply an inverse transform to the second magnitude spectrogram to create a second discrete signal that has a higher sampling rate than the first discrete audio signal.
G10L 25/18 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
G06N 3/088 - Non-supervised learning, e.g. competitive learning
G10L 19/02 - Speech or audio signal analysis-synthesis techniques for redundancy reduction, e.g. in vocodersCoding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
45.
Tokenization of text data to facilitate automated discovery of speech disfluencies
Introduced here are computer programs and associated computer-implemented techniques for discovering the presence of filler words through tokenization of a transcript derived from audio content. When audio content is obtained by a media production platform, the audio content can be converted into text content as part of a speech-to-text operation. The text content can then be tokenized and labeled using a Natural Language Processing (NLP) library. Tokenizing/labeling may be performed in accordance with a series of rules associated with filler words. At a high level, these rules may examine the text content (and associated tokens/labels) to determine whether patterns, relationships, verbatim, and context indicate that a term is a filler word. Any filler words that are discovered in the text content can be identified as such so that appropriate action(s) can be taken.
Introduced here are computer programs and associated computer-implemented techniques for discovering the presence of filler words through tokenization of a transcript derived from audio content. When audio content is obtained by a media production platform, the audio content can be converted into text content as part of a speech-to-text operation. The text content can then be tokenized and labeled using a Natural Language Processing (NLP) library. Tokenizing/labeling may be performed in accordance with a series of rules associated with filler words. At a high level, these rules may examine the text content (and associated tokens/labels) to determine whether patterns, relationships, verbatim, and context indicate that a term is a filler word. Any filler words that are discovered in the text content can be identified as such so that appropriate action(s) can be taken.
Introduced here are computer programs and associated computer-implemented techniques for facilitating the creation of a master transcription (or simply “transcript”) that more accurately reflects underlying audio by comparing multiple independently generated transcripts. The master transcript may be used to record and/or produce various forms of media content, as further discussed below. Thus, the technology described herein may be used to facilitate editing of text content, audio content, or video content. These computer programs may be supported by a media production platform that is able to generate the interfaces through which individuals (also referred to as “users”) can create, edit, or view media content. For example, a computer program may be embodied as a word processor that allows individuals to edit voice-based audio content by editing a master transcript, and vice versa.
Introduced here are computer programs and associated computer-implemented techniques for facilitating the creation of a master transcription (or simply “transcript”) that more accurately reflects underlying audio by comparing multiple independently generated transcripts. The master transcript may be used to record and/or produce various forms of media content, as further discussed below. Thus, the technology described herein may be used to facilitate editing of text content, audio content, or video content. These computer programs may be supported by a media production platform that is able to generate the interfaces through which individuals (also referred to as “users”) can create, edit, or view media content. For example, a computer program may be embodied as a word processor that allows individuals to edit voice-based audio content by editing a master transcript, and vice versa.
The invention relates to simultaneous recording and uploading systems and methods, and, more particularly to a simultaneous recording and uploading of multiple files from the same conversation.
The invention relates to audio drift normalization, and more particularly to audio drift normalization systems and methods that can normalize audio drift of a plurality of recordings from a source.
Different types of media experiences can be developed based on characteristics of the consumer. “Linear” experiences may require execution of a pre-built script, although the script could be dynamically modified by a media production platform. Linear experiences can include guided audio tours that are modified or updated based on the location of the consumer. “Enhanced” experiences include conventional media content that is supplemented with intelligent media content. For example, turn-by-turn directions could be supplemented with audio descriptions about the surrounding area. “Freeform” experiences, meanwhile, are those that can continually morph based on information gleaned from a consumer. For example, a radio station may modify what content is being presented based on the geographical metadata uploaded by a computing device associated with the consumer.
G06F 17/00 - Digital computing or data processing equipment or methods, specially adapted for specific functions
G06F 3/0484 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range
G10L 15/187 - Phonemic context, e.g. pronunciation rules, phonotactical constraints or phoneme n-grams
G06F 3/04817 - Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance using icons
G06F 16/68 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
G06F 16/683 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
41 - Education, entertainment, sporting and cultural services
Goods & Services
Providing business support services in the nature of start-up support for businesses of others related to podcasting and podcasting creation Audio recording and production services, namely, making podcasts and other audio content; providing a website featuring blogs in the field of podcasting, podcasting creation, and audio production
09 - Scientific and electric apparatus and instruments
35 - Advertising and business services
41 - Education, entertainment, sporting and cultural services
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable and recorded computer software platforms for making podcasts and other audio content; downloadable and recorded computer software platforms for recording, transcribing, editing, mixing podcasts and other media content; downloadable and recorded audio word processing computer software platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form Providing business support services related to podcasting and podcasting creation Audio recording and production services, namely, making podcasts and other audio content; providing a website featuring blogs in the field of podcasting, podcasting creation, and audio production Providing temporary use of on-line non-downloadable computer software for making podcasts and other audio content; platform as a service (PAAS) featuring computer software platforms for making podcasts and other audio content; providing temporary use of on-line non-downloadable computer software for recording, transcribing, editing, mixing podcasts and other media content; platform as a service (PAAS) featuring computer software platforms for recording, transcribing, editing, mixing podcasts and other media content; technical support services related to podcasting and podcasting creation, namely, troubleshooting in the nature of diagnosing computer software problems; providing temporary use of on-line non-downloadable computer software enabling editors and producers to edit sound files and writers to edit lyrics in text form; platform as a service (PAAS) featuring computer software and audio and word processing platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form
09 - Scientific and electric apparatus and instruments
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable computer software platforms for making podcasts and other audio content; downloadable computer software platforms for recording, transcribing, editing, and mixing podcasts and other media content; downloadable audio word processing computer software platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form Providing temporary use of on-line non-downloadable computer software for making podcasts and other audio content; platform as a service (PAAS) featuring computer software platforms for making podcasts and other audio content; providing temporary use of on-line non-downloadable computer software for recording, transcribing, editing, and mixing podcasts and other media content; platform as a service (PAAS) featuring computer software platforms for recording, transcribing, editing, and mixing podcasts and other media content; technical support services related to podcasting and podcasting creation, namely, troubleshooting in the nature of diagnosing computer software problems in the recording, creating, and editing of media; providing temporary use of on-line non-downloadable computer software enabling editors and producers to edit sound files and writers to edit lyrics in text form; platform as a service (PAAS) featuring computer software and audio and word processing platforms enabling editors and producers to edit sound files and writers to edit lyrics in text form
55.
Platform for producing and delivering media content
Media content can be created and/or modified using a network-accessible platform. Scripts for content-based experiences could be readily created using one or more interfaces generated by the network-accessible platform. For example, a script for a content-based experience could be created using an interface that permits triggers to be inserted directly into the script. Interface(s) may also allow different media formats to be easily aligned for post-processing. For example, a transcript and an audio file may be dynamically aligned so that the network-accessible platform can globally reflect changes made to either item. User feedback may also be presented directly on the interface(s) so that modifications can be made based on actual user experiences.
G11B 27/031 - Electronic editing of digitised analogue information signals, e.g. audio or video signals
G10L 19/008 - Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
Different types of media experiences can be developed based on characteristics of the consumer. “Linear” experiences may require execution of a pre-built script, although the script could be dynamically modified by a media production platform. Linear experiences can include guided audio tours that are modified or updated based on the location of the consumer. “Enhanced” experiences include conventional media content that is supplemented with intelligent media content. For example, turn-by-turn directions could be supplemented with audio descriptions about the surrounding area. “Freeform” experiences, meanwhile, are those that can continually morph based on information gleaned from a consumer. For example, a radio station may modify what content is being presented based on the geographical metadata uploaded by a computing device associated with the consumer.
G06F 3/0484 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range
G10L 15/187 - Phonemic context, e.g. pronunciation rules, phonotactical constraints or phoneme n-grams
G06F 3/0481 - Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance
G06F 16/68 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
G06F 16/683 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
57.
Platform for producing and delivering media content
Media content can be created and/or modified using a network-accessible platform. Scripts for content-based experiences could be readily created using one or more interfaces generated by the network-accessible platform. For example, a script for a content-based experience could be created using an interface that permits triggers to be inserted directly into the script. Interface(s) may also allow different media formats to be easily aligned for post-processing. For example, a transcript and an audio file may be dynamically aligned so that the network-accessible platform can globally reflect changes made to either item. User feedback may also be presented directly on the interface(s) so that modifications can be made based on actual user experiences.
G11B 27/031 - Electronic editing of digitised analogue information signals, e.g. audio or video signals
G10L 21/02 - Speech enhancement, e.g. noise reduction or echo cancellation
G10L 19/008 - Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing