Techniques are provided herein to provide a customized set of shader programs for different hardware versions for ray tracing hardware. Each such shader program varies in terms of what local state is stored. In particular, each shader program is tailored such that its associated set of local state does not exceed the amount actually needed by the corresponding hardware version. This customization of shader programs helps to reduce the amount of resources utilized by the shader programs, as compared with a situation in which a single “monolithic” shader program checked the hardware version at runtime, as such a monolithic shader program would need to reserve the maximum amount of resources for any possible hardware version.
Systems and methods for Neural Texture Block Compression (NTBC) of graphics textures using neural networks are disclosed. NTBC employs multi-layer perceptron (MLP) networks to map uncompressed textures to block-compressed formats, such as BC1 and BC4, achieving storage reductions while maintaining reasonable visual quality. The methods and systems use MLPs to compress textures in BC1 and BC4 formats. A first neural network comprises an endpoint network that predicts endpoints for each data point of a block of a texture. The second neural network is a color network which predicts original colors of uncompressed textures. To encode the inputs to these neural networks, multi-resolution feature grids are used that allows for usage of small MLPs and encourages model optimization.
H04N 19/186 - Procédés ou dispositions pour le codage, le décodage, la compression ou la décompression de signaux vidéo numériques utilisant le codage adaptatif caractérisés par l’unité de codage, c.-à-d. la partie structurelle ou sémantique du signal vidéo étant l’objet ou le sujet du codage adaptatif l’unité étant une couleur ou une composante de chrominance
G06T 7/90 - Détermination de caractéristiques de couleur
G06V 10/56 - Extraction de caractéristiques d’images ou de vidéos relative à la couleur
G06V 10/82 - Dispositions pour la reconnaissance ou la compréhension d’images ou de vidéos utilisant la reconnaissance de formes ou l’apprentissage automatique utilisant les réseaux neuronaux
H04N 19/176 - Procédés ou dispositions pour le codage, le décodage, la compression ou la décompression de signaux vidéo numériques utilisant le codage adaptatif caractérisés par l’unité de codage, c.-à-d. la partie structurelle ou sémantique du signal vidéo étant l’objet ou le sujet du codage adaptatif l’unité étant une zone de l'image, p. ex. un objet la zone étant un bloc, p. ex. un macrobloc
Techniques for performing machine learning operations are provided. The techniques include configuring a first portion of a first chiplet as a cache; performing caching operations via the first portion; configuring at least a first sub-portion of the first portion of the chiplet as directly-accessible memory; and performing machine learning operations with the first sub-portion by a machine learning accelerator within the first chiplet. In some examples, the first and second chiplets are separate dies. In some examples, the caching operations include one or more of storing a cache line evicted from a cache of the processing core or providing a cache line to the processing core in response to a miss in a cache of the processing core.
G06F 13/28 - Gestion de demandes d'interconnexion ou de transfert pour l'accès au bus d'entrée/sortie utilisant le transfert par rafale, p. ex. acces direct à la mémoire, vol de cycle
G06F 9/30 - Dispositions pour exécuter des instructions machines, p. ex. décodage d'instructions
G06F 9/38 - Exécution simultanée d'instructions, p. ex. pipeline ou lecture en mémoire
G06F 9/48 - Lancement de programmes Commutation de programmes, p. ex. par interruption
G06F 12/0893 - Mémoires cache caractérisées par leur organisation ou leur structure
G06F 12/128 - Commande de remplacement utilisant des algorithmes de remplacement adaptée aux systèmes de mémoires cache multidimensionnelles, p. ex. associatives d’ensemble, à plusieurs mémoires cache, multi-ensembles ou multi-niveaux
G06F 18/214 - Génération de motifs d'entraînementProcédés de Bootstrapping, p. ex. ”bagging” ou ”boosting”
In an implementation, a flip-flop circuit may include a master latch configured to receive and store input data. The flip-flop circuit may also include a slave latch coupled to the master latch and a clock gating circuit coupled between the master latch and the slave latch, where the clock gating circuit is configured to selectively enable or disable an internal gated clock signal to the slave latch based on a state of the master latch.
A system can include an integrated circuit; and an adaptive voltage regulator (AVR) device configured to provide a power supply signal having a regulated voltage to the integrated circuit. The AVR can include a slew rate control circuit configured to perform operations including adjusting a maximum slew rate of the regulated voltage based on an amplitude of a load current provided to the integrated circuit (IC). In some examples, the AVR improves the performance of the IC by responding rapidly to changes in the IC's target supply voltage when the load current is relatively low, while also limiting the occurrence of potentially harmful over-currents by lowering the slew rate of the regulated supply voltage when the load current is already relatively high.
G05F 1/575 - Régulation de la tension ou de l'intensité là où la variable effectivement régulée par le dispositif de réglage final est du type continu utilisant des dispositifs à semi-conducteurs en série avec la charge comme dispositifs de réglage final caractérisé par le circuit de rétroaction
G05F 1/565 - Régulation de la tension ou de l'intensité là où la variable effectivement régulée par le dispositif de réglage final est du type continu utilisant des dispositifs à semi-conducteurs en série avec la charge comme dispositifs de réglage final sensible à une condition du système ou de sa charge en plus des moyens sensibles aux écarts de la sortie du système, p. ex. courant, tension, facteur de puissance
Systems, methods, and devices for rendering an image. A model of a three-dimensional (3D) object is rendered to generate a two-dimensional (2D) image. A shading rate of the rendering is based on GPU occupancy and user interaction. In some implementations, the GPU occupancy and user interaction comprise GPU occupancy and user interaction during rendering of a prior frame. In some implementations, the GPU occupancy and user interaction comprise GPU occupancy and user interaction during rendering of an immediately preceding frame. In some implementations, the GPU occupancy and user interaction comprise GPU occupancy and user interaction during rendering of a plurality of prior frames. In some implementations, the GPU occupancy and user interaction comprise average GPU occupancy and average user interaction over rendering of a plurality of prior frames.
An apparatus and method for performing efficient dynamic scheduling of kernels in a processing circuit. In various implementations, a host processing circuit conveys to an accelerator device, an input vector and a representation of a sparse matrix for a matrix multiplication operation. To generate the output vector for the matrix multiplication operation, multiple compute circuits of the accelerator device overlap execution of multiple types of operations. Four types of operations include classifying operations that classify rows of the sparse matrix as short rows or long rows, matrix multiplication operations using long rows, partitioning operations that group short rows together for subsequent matrix multiplication operations, and matrix multiplication operations using short rows. To overlap the operations, the multiple compute circuits generate kernels without involvement from the host processing circuit. The multiple compute circuits convey the generated kernels to a scheduling queue of a control circuit of the accelerator device.
A semiconductor device includes one or more active devices disposed between a processor die and a package substrate. The semiconductor device includes a first layer with a processor die, a second layer with one or more active devices, and a third layer with a package substrate, where the second layer is disposed between the first and third layers. The one or more active devices are semiconductor-based devices, such as voltage regulators, that participate in supplying power to the processor die and are electrically connected to the processor die using various connection configurations. The implementations use short path lengths for improved performance with a compact structure that avoids the use of edge wiring or interposers without occupying processor die space. Implementations include the use of through-silicon vias (TSVs) to provide short path lengths while reducing the number of connection resources used by the one or more power components.
Disclosed herein are chip packages and electronic devices that utilized a silicon bridge having a memory controller to interface between a logic device having at least one compute die and one or more memory stacks within a singular chip package. In one example, a chip package is provided that a substrate, a plurality of silicon bridges, a redistribution layer, a logic device, and a memory stack. The plurality of silicon bridges are disposed in a common layer and electrically and mechanically coupled to the substrate. At least a first silicon bridge of the plurality of silicon bridges includes a plurality of decoupling capacitors. The redistribution layer is disposed on the plurality of silicon bridges. The logic device is disposed over the redistribution layer and includes one or more compute dies. The memory stack is disposed over the redistribution layer adjacent the logic device.
An apparatus and method for efficiently performing efficient data storage and data transfer of machine learning (ML) data. In various implementations, a host processing circuit of a computing system executes an ML application. The application includes a computational graph that provides the computational order of the ML nodes, layers and stages of the ML model. The host processing circuit translates function calls in the application to commands particular to an accelerator circuit. When data storage capacity of the local memory of the accelerator circuit exceeds a threshold, the accelerator circuit compress one or more of the weights and intermediate data prior to offloading this data from the local memory based on completion of execution of a corresponding ML layer and a compression ratio being greater than a ratio based on decompression throughput and data transfer rates between levels of a memory hierarchy.
This disclosure describes techniques for generating oriented bounding boxes within a bounding volume hierarchy. Use of OBBs speeds up ray tracing by reducing false positive hits. However, computing the orientation for each box is typically a complex process. According to this disclosure, a BVH builder generates a BVH having oriented bounding boxes by generating bounding boxes and determining an orientation for one or more such bounding boxes. To generate the orientation for a particular node having a bounding box, the BVH builder selects the orientation of one of the children of the node and uses that orientation as the orientation for the node. This selection is performed without consideration of any other child. Rather than attempting to find the child with the “best” orientation, or finding an orientation that accounts for the orientation of each child, the BVH builder considers the orientation of only one child, which is computationally cheaper.
Techniques are provided for dynamically adjusting memory states in response to a density of received Column Address Strobe (CAS) memory commands. A memory controller monitors CAS command density over a configurable quantity of memory cycles in an idle memory state. If the CAS command density exceeds an idling threshold during the idle memory state, the memory subsystem selectively throttles CAS commands for a defined throttling duration to limit bandwidth and stabilize power demands, after which the memory subsystem transitions to a normal memory state in which all CAS commands are processed. If CAS command density subsequently falls below the idling threshold, the memory subsystem returns to the idle state.
Methods, systems, and computer-readable media for runtime management using accelerator-resident managers. A first manager resident in a first accelerator accesses a runtime representation stored in shared memory accessible by resident managers of a plurality of accelerators. The runtime representation identifies dependencies among kernels of an application and assignments of the kernels to Accelerated Processing Units (APUs) of the plurality of accelerators. The first manager launches kernels assigned to APUs of the first accelerator and updates the runtime representation based on availability information stored in the shared memory for APUs of the first accelerator and APUs of a second accelerator. The first manager updates the runtime representation to change an assignment of at least one kernel from an APU of the second accelerator to an APU of the first accelerator for a later execution iteration.
A voltage regulator assembly can include a power stage chip and an inductive component coupled to the power stage chip. A footprint of the power stage chip can be bonded to a surface of a printed circuit board. The inductive component can include at least one coil, wherein a first portion of the coil is disposed on or over a surface of the power stage chip. The printed circuit board, the power stage chip, and the first portion of the coil can be arranged in a stack. The voltage regulator assembly can be configured to provide one or more power supply signals via the inductive component.
Disclosed herein are integrated circuit (IC) dies, and electronic devices including the same, that include flexible routing arrangements at a hybrid bonding layer featuring an etch stop layer. The etch stop layer enables vias connected with the hybrid bonding layer to terminate at both contact pads formed on the surface of the IC die and buried conductive pads formed on a top metal layer of the IC die.
H01L 23/538 - Dispositions pour conduire le courant électrique à l'intérieur du dispositif pendant son fonctionnement, d'un composant à un autre la structure d'interconnexion entre une pluralité de puces semi-conductrices se trouvant au-dessus ou à l'intérieur de substrats isolants
H01L 21/768 - Fixation d'interconnexions servant à conduire le courant entre des composants distincts à l'intérieur du dispositif
H01L 25/18 - Ensembles consistant en une pluralité de dispositifs à semi-conducteurs ou d'autres dispositifs à l'état solide les dispositifs étant de types prévus dans plusieurs différents groupes principaux de la même sous-classe , , , , ou
16.
EFFICIENT BATCHING FOR BOUNDING VOLUME HIERARCHY BUILD OPERATIONS
Ray tracing workloads utilize bounding volume hierarchies as an acceleration structure for efficiency. These structures are built for scene geometry. The build operations are batched together and executed in parallel. Such operations are performed in a set of phases separated by synchronization barriers. An important issue with such batching is that BVH build operations have a wide range of execution times. Due to the barriers, slower BVH builds can cause delays in BVH builds that would otherwise complete more quickly. For this reason, instead of batching build requests together regardless of how long they might take to execute, a driver sorts BVH build requests into bins based on a sorting criterion which acts as a proxy for execution time. Because barriers only apply to bins they are in, barriers apply for work that is estimated to complete in a similar amount of time, alleviating the above issue.
Accelerated circuit design test coverage closure includes generating, by computer hardware, constraint ranges for input variables for a circuit design based on constraints for the input variables. A first value set for a verification testbench for the circuit design is generated. The first value set includes randomly generated values for the input variables constrained based on the constraint ranges. Whether the first value set is unique is detected by comparing the first value set with training value sets of a training corpus. In response to detecting that the first value set is unique, the first value is classified set by assigning a label selected from a plurality of labels to the first value set and adding the first value set to the training corpus.
Systems and methods for generating aligned data from segmented data channels include a memory controller with one or more buffers. The memory controller stores mismatched, segmented chunks of data received from a memory in buffers such that, once a complete set of chunks for a value requested from memory is received, the memory controller outputs all of the chunks substantially simultaneously as aligned data. Various data output priorities for the memory controller include an oldest-first order, a round robin order, and an order based on a priority or tag associated with chunks of data.
An apparatus and method for efficiently performing efficient data storage and data transfer of machine learning (ML) data. In various implementations, a host processing circuit of a computing system executes an ML application. The application includes a computational graph that provides the computational order of the ML nodes, layers and stages of the ML model. The host processing circuit translates function calls in the application to commands particular to an accelerator circuit. When data storage capacity of the local memory of the accelerator circuit exceeds a threshold, the accelerator circuit compress one or more of the weights and intermediate data prior to offloading this data from the local memory based on completion of execution of a corresponding ML layer and a compression ratio being greater than a ratio based on decompression throughput and data transfer rates between levels of a memory hierarchy.
An electronic device includes a memory device that is able to block execution of commands. The memory device includes memory circuitry and control circuitry. The memory circuitry stores data. The control circuitry is connected to the memory circuitry. The control circuitry receives a first signal from a host device, and blocks execution of commands to the memory circuitry based on a parity error within the first signal.
An electronic device includes a memory device that is able to block execution of commands. The memory device includes memory circuitry and control circuitry. The memory circuitry stores data. The control circuitry is connected to the memory circuitry. The control circuitry receives a first signal from a host device, and blocks execution of commands to the memory circuitry based on a parity error within the first signal.
G06F 3/06 - Entrée numérique à partir de, ou sortie numérique vers des supports d'enregistrement
G06F 11/10 - Détection ou correction d'erreur par introduction de redondance dans la représentation des données, p. ex. en utilisant des codes de contrôle en ajoutant des chiffres binaires ou des symboles particuliers aux données exprimées suivant un code, p. ex. contrôle de parité, exclusion des 9 ou des 11
22.
DIRECT BROADCASTING OF CALCULATED SUMS BETWEEN ACCUMULATORS IN A MULTI-CORE SPATIAL ARCHITECTURE
An accumulator broadcast network is described that permits an accumulator in one core of a multi-core IC to transmit high-precision data to accumulators in multiple cores in parallel. That is, instead of a sending data to local memory and the broadcasting to different cores using shared memory or messaging protocols, the data stored in the accumulator registers can be directly transmitted to other accumulators.
G06F 15/80 - Architectures de calculateurs universels à programmes enregistrés comprenant un ensemble d'unités de traitement à commande commune, p. ex. plusieurs processeurs de données à instruction unique
An integrated circuit (IC) module having plurality of devices embedded into a substrate core is provided, along with electronic devices and chip packages having the same. Also provided are method for fabricating the IC module. In one example, the IC module includes a substrate core, an interposer, a plurality of devices, and a molded encapsulation material. The substrate core includes a first face, a second face, and a cavity formed in the substrate core. The cavity is open through the first and second faces of the substrate core. The interposer is disposed in the cavity and has a first face oriented in a common direction with the first surface. The plurality of devices are disposed in the cavity and mounted on a second face of the interposer. The molded encapsulation material is disposed in the cavity and encapsulates the plurality of devices against the second face of the interposer.
H01L 23/31 - Encapsulations, p. ex. couches d’encapsulation, revêtements caractérisées par leur disposition
H01L 21/56 - Encapsulations, p. ex. couches d’encapsulation, revêtements
H01L 23/18 - Matériaux de remplissage caractérisés par le matériau ou par ses propriétes physiques ou chimiques, ou par sa disposition à l'intérieur du dispositif complet
H01L 23/367 - Refroidissement facilité par la forme du dispositif
H01L 23/498 - Connexions électriques sur des substrats isolants
24.
CLOCK STRETCH COMPENSATION FOR DIGITAL FREQUENCY-LOCKED LOOP CIRCUITS
An implementation is a processor including a voltage regulator configured to provide a supply voltage, a digital frequency-locked loop (DFLL) circuit having a stretch response configured to reduce an operating frequency in response to voltage droops in the supply voltage, and a clock stretch compensation (CSC) circuit. The CSC circuit being configured to monitor the operating frequency relative to a target frequency at a predetermined sampling rate, and increase the supply voltage based on a difference between the target frequency and the operating frequency.
H03K 5/06 - Mise en forme d'impulsions par augmentation de duréeMise en forme d'impulsions par diminution de durée par l'utilisation de lignes à retard ou d'autres éléments à retard analogues
H03K 5/135 - Dispositions ayant une sortie unique et transformant les signaux d'entrée en impulsions délivrées à des intervalles de temps désirés par l'utilisation de signaux de référence de temps, p. ex. des signaux d'horloge
H03K 5/19 - Contrôle de la configuration de trains d'impulsions
A processor includes a scalable matrix arithmetic unit with multiple block scale registers to provide scale data to different groups of arithmetic elements. The matrix arithmetic unit employs the arithmetic elements to perform matrix operations, such as outer product operations, using provided arithmetic operands. The arithmetic elements are configured to apply the scale data to the operands prior to or during the matrix operations. By employing block scale registers to provide the scale data to the groups of arithmetic elements, the processor is able to efficiently implement scale operations without consuming an undesirably large amount of circuit area.
Embodiments herein relate to implementing a dynamic scheduler within a heterogeneous computing system. At runtime, there may be various circumstances the computing system is faced with that makes the usage of one type of processor more optimum than another processor for a certain number or type of calculations. When presented with an inference task, at runtime, a dynamic scheduler can determine one set of calculations that should be made on one processor, and another set of calculations that should be made on another, different processor. The separate calculations can then be combined to produce the result for the inference task.
A processor includes a scalable matrix arithmetic unit with multiple block scale registers to provide scale data to different groups of arithmetic elements. The matrix arithmetic unit employs the arithmetic elements to perform matrix operations, such as outer product operations, using provided arithmetic operands. The arithmetic elements are configured to apply the scale data to the operands prior to or during the matrix operations. By employing block scale registers to provide the scale data to the groups of arithmetic elements, the processor is able to efficiently implement scale operations without consuming an undesirably large amount of circuit area.
42 - Services scientifiques, technologiques et industriels, recherche et conception
Produits et services
Integrated circuit design services; computer chip and processor design services; computer hardware design and development consulting; Research and development services, namely, providing research development services for others and information on new products in the fields of semiconductors, integrated circuits, memory devices and computer hardware and software; computer hardware and software consulting services; Software as a service (SAAS) featuring software development tools for the creation of software applications and application interfaces, Software as a service (SAAS) services featuring software for use in accelerated computing, high performance computing (HPC) and machine learning (ML) computing; Updating of computer software; Installation of computer software; Maintenance of computer software; Computer programming; Computer software design; design of computer software in the field of semiconductors, integrated circuits, and memory devices; Computer hardware and firmware design consultancy; Development of computer chip design software; rental of computer software; data conversion of computer programs and data, not physical conversion; Providing online non-downloadable computer software for building and optimizing AI, machine learning (ML), and high performance computing (HPC) workloads on computer hardware; Providing online non-downloadable computer software for cloud computing; Consultancy services in the fields of cloud computing and open source software for use in software development; maintenance and updating of cloud-based design and development software tools
29.
REFRESH COMMAND FOR MULTIPLE MEMORY BANKS OF MULTIPLE MEMORY DIES
A memory system includes one or more memory dies and a memory controller. The one or more memory dies have a first memory bank, a second memory bank, a third memory bank, and a fourth memory bank. The first memory bank is associated with a first identifier and the third memory bank is associated with a second identifier. The memory controller circuitry is coupled to the one or more memory dies. The memory controller circuitry outputs a refresh command to the one or more memory dies to refresh the first memory bank and the third memory bank during a first period. The second memory bank and the fourth memory bank are accessible by the memory controller circuitry during the first period.
A processing system includes an automation circuit configured to extract, independent of one or more application programming interfaces associated with a computational environment executing at the processing system, application state data representing a current state of the computational environment. The automation circuit is further configured to generate one or more objectives based on application state data representing a current state of a computational environment executing at the processor, and decompose the one or more objectives into a plurality of sub-objectives. The automation circuit is also configured to generate executable code corresponding to the plurality of sub-objectives, and execute the executable code to perform one or more actions in the computational environment.
A63F 13/60 - Création ou modification du contenu du jeu avant ou pendant l’exécution du programme de jeu, p. ex. au moyen d’outils spécialement adaptés au développement du jeu ou d’un éditeur de niveau intégré au jeu
A graphics processing unit (GPU) schedules recurrent matrix multiplication operations at different subsets of CUs of the GPU. The GPU includes a scheduler that receives sets of recurrent matrix multiplication operations, such as multiplication operations associated with a recurrent neural network (RNN). The multiple operations associated with, for example, an RNN layer are fused into a single kernel, which is scheduled by the scheduler such that one work group is assigned per compute unit, thus assigning different ones of the recurrent matrix multiplication operations to different subsets of the CUs of the GPU. In addition, via software synchronization of the different workgroups, the GPU pipelines the assigned matrix multiplication operations so that each subset of CUs provides corresponding multiplication results to a different subset, and so that each subset of CUs executes at least a portion of the multiplication operations concurrently.
To convert a value from an initial data format to a modified data format having a modified data format style, an accelerator unit executes a first instruction to convert the value from the initial data format to an intermediate data format in an intermediate data format style. The accelerator unit then rounds the value in the intermediate data format according to an intermediate rounding mode. After rounding the value, the accelerator unit executes a second instruction to convert the value to the modified data format in the modified data format style. Further, the accelerator unit rounds the value according to a modified rounding mode such that converting the value from the initial data format to the modified data format introduces a single rounding error.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06F 9/30 - Dispositions pour exécuter des instructions machines, p. ex. décodage d'instructions
G06F 7/499 - Maniement de valeur ou d'exception, p. ex. arrondi ou dépassement
33.
PROCESSING UNIT CONFIGURED TO CONVERT THE DATA FORMAT OF DATA ELEMENT VALUES USING AN INTERMEDIATE DATA FORMAT
To convert a value from an initial data format to a modified data format having a modified data format style, an accelerator unit executes a first instruction to convert the value from the initial data format to an intermediate data format in an intermediate data format style. The accelerator unit then rounds the value in the intermediate data format according to an intermediate rounding mode. After rounding the value, the accelerator unit executes a second instruction to convert the value to the modified data format in the modified data format style. Further, the accelerator unit rounds the value according to a modified rounding mode such that converting the value from the initial data format to the modified data format introduces a single rounding error.
In an implementation, a clock latch system may include a front-end circuit configured to receive and process input signals. The clock latch system may also include a latch circuit connected to the front-end circuit and configured to store a state of a processed signal. The clock latch system may also include an output driver circuit connected to the latch circuit. The clock latch system may also include test enable signals coupled to at least one of the latch circuit or the output driver circuit. The clock latch system may also include a clock generation circuit connected to the latch circuit.
H03K 19/20 - Circuits logiques, c.-à-d. ayant au moins deux entrées agissant sur une sortieCircuits d'inversion caractérisés par la fonction logique, p. ex. circuits ET, OU, NI, NON
A scale register is configured to store a first set of block scale factors associated with a first input vector to an operation and a second set of block scale factors associated with a second input vector to the operation. A processing unit is configured to perform the operation based on the first input vector, a first scale factor selected from the first set based on user input, the second input vector, and a second scale factor selected from the second set based on the user input. In some cases, the operation is an outer product operation and the processing unit includes a matrix engine configured to perform the outer product. The elements of the first input vector and the second input vector can represent quantized values of elements represented in a first precision that is higher than a second precision of the quantized values.
G06F 7/483 - Calculs avec des nombres représentés par une combinaison non linéaire de nombres codés, p. ex. nombres rationnels, système de numération logarithmique ou nombres à virgule flottante
Systems, apparatuses, and methods for reducing memory power consumption without substantial performance impact by selectively delaying non-critical memory requests are disclosed. A system management unit transfers an amount of power allocated from a memory subsystem to other component(s) responsive to detecting a first condition. In one embodiment, the first condition is detecting one or more processors having tasks to execute. In response to the system management unit transferring the amount of power from the memory subsystem to one or more processors, a memory controller delays non-critical memory requests while performing critical memory requests to memory.
An accelerator unit (AU) includes tile registers configured to store matrices associated with a machine-learning model to be implemented. Further, the AU supports a tile read column instruction that reads a column of elements from a matrix stored in the tile registers. When executing a tile read column instruction for a matrix, the AU first determines whether a transposed clone of the matrix is stored in the tile registers. If there is no transposed clone stored the AU generates a transposed clone by performing row read and tile write commands using the matrix. After confirming a transposed clone of the matrix is in the tile registers, the AU reads the rows of the transposed clone corresponding to the column indicated in the tile read column instruction.
G06F 9/50 - Allocation de ressources, p. ex. de l'unité centrale de traitement [UCT]
G06F 7/78 - Dispositions pour le réagencement, la permutation ou la sélection de données selon des règles prédéterminées, indépendamment du contenu des données pour changer l'ordre du flux des données, p. ex. transposition matricielle ou tampons du type pile d'assiettes [LIFO]Gestion des occurrences du dépassement de la capacité du système ou de sa sous-alimentation à cet effet
A scale register is configured to store a first set of block scale factors associated with a first input vector to an operation and a second set of block scale factors associated with a second input vector to the operation. A processing unit is configured to perform the operation based on the first input vector, a first scale factor selected from the first set based on user input, the second input vector, and a second scale factor selected from the second set based on the user input. In some cases, the operation is an outer product operation and the processing unit includes a matrix engine configured to perform the outer product. The elements of the first input vector and the second input vector can represent quantized values of elements represented in a first precision that is higher than a second precision of the quantized values.
A process of making a stacked semiconductor die can include a step of forming a conformal film of insulating material on one or more surfaces of a singulated semiconductor die. In some examples, the presence of the film tends to reduce the formation of cracks during the die stacking process.
H01L 23/31 - Encapsulations, p. ex. couches d’encapsulation, revêtements caractérisées par leur disposition
C23C 16/50 - Revêtement chimique par décomposition de composés gazeux, ne laissant pas de produits de réaction du matériau de la surface dans le revêtement, c.-à-d. procédés de dépôt chimique en phase vapeur [CVD] caractérisé par le procédé de revêtement au moyen de décharges électriques
H01L 21/56 - Encapsulations, p. ex. couches d’encapsulation, revêtements
H01L 23/00 - Détails de dispositifs à semi-conducteurs ou d'autres dispositifs à l'état solide
Synthetic wafer generation based on a multi-tiered ML model that includes a technology-level ML model that infers coordinate-based wafer-level characteristics of a synthetic wafer model, and a product-level ML model that infers coordinate-based die-level characteristics of the synthetic wafer model based on the inferred coordinate-based wafer-level characteristics, a design of dies of the synthetic wafer model, and coordinates of the dies. The multi-tiered ML model may be trained based on sparse training data using a combinations of supervised learning and generative learning methods, which may include a tabular variational autoencoder using metric learning (TVAE) method, a conditional tabular generative adversarial network (CTGAN) method, Gaussian Copulas, and/or K-nearest neighbor method. Wafer-level training characteristics may be converted from a coordinate system of a test probe to wafer coordinates (e.g., polar coordinates).
G06F 30/27 - Optimisation, vérification ou simulation de l’objet conçu utilisant l’apprentissage automatique, p. ex. l’intelligence artificielle, les réseaux neuronaux, les machines à support de vecteur [MSV] ou l’apprentissage d’un modèle
G06F 119/02 - Analyse de fiabilité ou optimisation de fiabilitéAnalyse de défaillance, p. ex. performance dans le pire scénario, analyse du mode de défaillance et de ses effets [FMEA]
41.
LIGHT INCREMENTAL FLOW FOR QUALITY-OF-RESULT RESILIENCE
Implementing a circuit design for a target integrated circuit includes receiving guide information including placement information for a circuit design from a prior implementation flow. A light incremental flow is performed on the circuit design using the guide information. The light incremental flow generates a placed and routed circuit design that re-uses at least a portion of the placement information from the prior implementation flow.
A system and method mitigates unauthorized access by a network interface controller. A transaction associated with a target resource within the network interface controller is received by the network interface controller. Further, the transaction is authorized based on a characteristic of the transaction and a characteristic of the target resource within the network interface controller. The transaction is output to the target resource based on the transaction being authorized.
An integrated circuit (IC) device package includes a PCB circuitry having a back side opposing a topside, a package substrate having a die side and a landside that is electrically connected to the topside of the PCB circuitry, a die electrically connected to the die side of the package substrate; and a landside capacitor, the landside capacitor electrically connected to both the landside of the package substrate and the topside of the PCB circuitry
Embodiments herein describe a pull-based model to dispatch tasks in an accelerator device. That is, rather than a push-based model where a connected host pushes tasks into hardware (HW) queues in the accelerator device, the embodiments herein describe a pull-based model where a command processor (CP) loads tasks into the HW queues after any data dependencies have been resolved.
Devices, methods and systems for managing resources in a computing device. Information regarding resource usage is captured. A prediction is generated, based on the information, that resource usage by a processor will exceed a threshold during an upcoming time. An operating parameter of the processor is adjusted, based on the prediction. In some implementations, information regarding memory bandwidth is captured. A prediction is generated, based on the information, that a memory region stored in a first memory device will be addressed by a memory intensive instruction during an upcoming time period. Data stored in the memory region is moved to a second memory device, based on the prediction.
A processor supports selective provision and blocking of content (e.g., video or other visual content) to different display devices based on the encryption standards of those display devices. The processor identifies the encryption standard required for the content (that is, the process by which the content has been or is to be encrypted) and whether the encryption standard for each display device meets the encryption standard of the content. The processor then provides the content to the display devices whose encryption standards meet the standard of the content, and blocks provision of the content to any display device that does not meet the encryption standard of the content.
H04N 21/2347 - Traitement de flux vidéo élémentaires, p. ex. raccordement de flux vidéo ou transformation de graphes de scènes du flux vidéo codé impliquant le cryptage de flux vidéo
H04N 21/2343 - Traitement de flux vidéo élémentaires, p. ex. raccordement de flux vidéo ou transformation de graphes de scènes du flux vidéo codé impliquant des opérations de reformatage de signaux vidéo pour la distribution ou la mise en conformité avec les requêtes des utilisateurs finaux ou les exigences des dispositifs des utilisateurs finaux
A processor protects a machine learning model (MLM) from unauthorized access. The processor employs a neural processing unit (NPU) to execute the MLM and implements decryption and encryption processes to decrypt the MLM and re-encrypt the MLM at different points along MLM storage and execution paths. Furthermore, the processor executes the encryption and decryption processes at different processing units and processing engines, thereby reducing the ability of malicious software to access the MLM. In addition, the processor protects buffers of the NPU from unauthorized access.
G06F 21/62 - Protection de l’accès à des données via une plate-forme, p. ex. par clés ou règles de contrôle de l’accès
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
48.
SYSTEMS AND METHODS FOR EFFICIENT EXECUTION OF VECTOR OPERATIONS
A disclosed computer-implemented method may include loading a pair of input vectors of into a respective pair of registers included in a processor, each input vector associated with a different shared scale term. The method may also include storing a shared scale term corresponding to at least one of the pair of input vectors within at least one control register of the processor. The method may also include performing a vector operation that utilizes the pair of input vectors by accessing the at least one control register to retrieve the shared scale term as part of the vector operation Various other methods, systems, and computer-readable media are also disclosed.
G06F 9/30 - Dispositions pour exécuter des instructions machines, p. ex. décodage d'instructions
G06F 7/483 - Calculs avec des nombres représentés par une combinaison non linéaire de nombres codés, p. ex. nombres rationnels, système de numération logarithmique ou nombres à virgule flottante
Disclosed devices, systems, and methods may enhance communication protocols for low latency applications. Systems may include a device and a host interconnected by a high-performance interconnect and/or communication link. The device may comprise a transmitter, a receiver, and a control unit that may manage a credit-based flow control mechanism. In some aspects, the device may initiate a push write request, send a data header with an identifier (UQID) matching the push write request identifier (CQID), and transmit the data payload. The host may receive the push write request, match the UQID with the CQID, perform the write operation, and send a completion message back to the device. The method may involve ensuring sufficient credits before initiating the push write transaction, which may help prevent data loss and ensure reliable delivery. The push write mechanism may reduce the number of link traversals required for device-to-host memory writes, potentially lowering overall latency.
Systems and methods for correlating performance on bare metal systems with virtualized instances such as those commonly used in cloud computing systems are disclosed. Performance data may be collected from different bare metal and cloud instances. The performance data may be used to predict the performance of an application on another system, even if the particular performance counter of interest is unavailable on the system. Using measured and estimated performance counters (e.g., instructions counters), multiple measured and unmeasured but predicted instances can be compared and sorted to assist the user in making an informed decision when selecting where to run their instance and what configuration to use.
H04L 41/0806 - Réglages de configuration pour la configuration initiale ou l’approvisionnement, p. ex. prêt à l’emploi [plug-and-play]
H04L 41/50 - Gestion des services réseau, p. ex. en assurant une bonne réalisation du service conformément aux accords
H04L 41/5054 - Déploiement automatique des services déclenchés par le gestionnaire de service, p. ex. la mise en œuvre du service par configuration automatique des composants réseau
H04L 67/10 - Protocoles dans lesquels une application est distribuée parmi les nœuds du réseau
51.
LOW SWING REPEATER CIRCUIT WITH TRANSITION CONTROL
In an implementation, a repeater circuit includes an input circuit configured to receive an input signal, a control circuit coupled to the input circuit, an output circuit coupled to the control circuit and configured to drive an output signal, and a feedback circuit coupled between the output circuit and the control circuit, wherein the control circuit is configured to operate the output circuit in a switching mode and operate the output circuit in a diode-connected mode to provide a reduced voltage swing responsive to the feedback circuit.
A data processing system includes a memory accessing agent and a memory controller coupled to the memory accessing agent. The memory controller includes an ECC check circuit for detecting errors in a data element and extracting metadata from an error correcting code, in which the detecting and extracting includes forming a plurality of error statuses based on the data element and the error correcting code for different combinations of metadata, and picking a final status and final metadata based on the plurality of error statuses.
G06F 11/10 - Détection ou correction d'erreur par introduction de redondance dans la représentation des données, p. ex. en utilisant des codes de contrôle en ajoutant des chiffres binaires ou des symboles particuliers aux données exprimées suivant un code, p. ex. contrôle de parité, exclusion des 9 ou des 11
A memory controller system comprises a first memory controller for accessing a first sub-channel of a memory module, and a second memory controller for accessing a second sub-channel of the memory module. Each of the first and second memory controllers is configurable as one of a master memory controller and a slave memory controller. The master memory controller is enabled to send a global command to the memory module in response to a request, and the slave memory controller is disabled from sending the global command to the memory module in response to the request.
G11C 11/406 - Organisation ou commande des cycles de rafraîchissement ou de régénération de la charge
G11C 11/4096 - Circuits de commande ou de gestion d'entrée/sortie [E/S, I/O] de données, p. ex. circuits pour la lecture ou l'écriture, circuits d'attaque d'entrée/sortie ou commutateurs de lignes de bits
54.
MACHINE LEARNING-BASED OPTIMIZATION OF PROCESSING SYSTEM HARDWARE
A processing system includes a plurality of hardware components and an intelligent optimization component. The intelligent optimization component includes at least one inference engine configured to receive a first input comprising a first encoded representation of at least a first hardware component of the plurality of hardware components and a second input including a second encoded representation of an application. Based on these inputs, the inference engine generates an output including at least one recommended hardware component upgrade for the processing system. The intelligent optimization component presents the at least one recommended hardware component upgrade to a user on a display coupled to the processing system.
G06F 30/27 - Optimisation, vérification ou simulation de l’objet conçu utilisant l’apprentissage automatique, p. ex. l’intelligence artificielle, les réseaux neuronaux, les machines à support de vecteur [MSV] ou l’apprentissage d’un modèle
In an implementation, a self-gating system includes a front-end circuit configured to generate an enable output based on input data signals corresponding to multiple flip-flop bits, a latch circuit coupled to the front-end circuit and configured to receive the enable output, an output driver circuit coupled to the latch circuit and configured to generate a gated clock signal, and a flip-flop circuit coupled to the output driver circuit and configured to receive the gated clock signal, wherein the flip-flop circuit includes multiple flip-flops corresponding to the multiple flip-flop bits.
A processor supports selective provision and blocking of content (e.g., video or other visual content) to different display devices based on the encryption standards of those display devices. The processor identifies the encryption standard required for the content (that is, the process by which the content has been or is to be encrypted) and whether the encryption standard for each display device meets the encryption standard of the content. The processor then provides the content to the display devices whose encryption standards meet the standard of the content, and blocks provision of the content to any display device that does not meet the encryption standard of the content.
A disclosed computer-implemented method may include loading a pair of input vectors of into a respective pair of registers included in a processor, each input vector associated with a different shared scale term. The method may also include storing a shared scale term corresponding to at least one of the pair of input vectors within at least one control register of the processor. The method may also include performing a vector operation that utilizes the pair of input vectors by accessing the at least one control register to retrieve the shared scale term as part of the vector operation Various other methods, systems, and computer-readable media are also disclosed.
Disclosed devices, systems, and methods may enhance communication protocols for low latency applications. Systems may include a device and a host interconnected by a high-performance interconnect and/or communication link. The device may comprise a transmitter, a receiver, and a control unit that may manage a credit-based flow control mechanism. In some aspects, the device may initiate a push write request, send a data header with an identifier (UQID) matching the push write request identifier (CQID), and transmit the data payload. The host may receive the push write request, match the UQID with the CQID, perform the write operation, and send a completion message back to the device. The method may involve ensuring sufficient credits before initiating the push write transaction, which may help prevent data loss and ensure reliable delivery. The push write mechanism may reduce the number of link traversals required for device-to-host memory writes, potentially lowering overall latency.
An apparatus includes first and second circuitry, such as a first memory and a second memory. The first circuitry is configured to store compressed information representing a set of distribution functions associated with a set of entropy models. The second circuitry is configured to store mappings of requests for subsets of the set of distribution functions to one or more offsets in the first circuitry corresponding to the compressed information representing the subsets. The apparatus also includes circuitry that implements the set of entropy models. This circuitry is configured to perform entropy coding based on the subsets and a first entropy model selected from the set of entropy models. This circuitry can include one or more entropy engines that have a set of ports for encoding or decoding one or more streams of bits or symbols. Multiple streams can therefore be encoded or decoded concurrently or in parallel.
H04N 19/13 - Codage entropique adaptatif, p. ex. codage adaptatif à longueur variable [CALV] ou codage arithmétique binaire adaptatif en fonction du contexte [CABAC]
H04N 19/423 - Procédés ou dispositions pour le codage, le décodage, la compression ou la décompression de signaux vidéo numériques caractérisés par les détails de mise en œuvre ou le matériel spécialement adapté à la compression ou à la décompression vidéo, p. ex. la mise en œuvre de logiciels spécialisés caractérisés par les dispositions des mémoires
60.
LOOKUP TABLE ACCESS FOR QUANTIZATION AND DEQUANTIZATION OPERATIONS
One or more cache lines are configured to store a lookup table (LUT) having a plurality of floating-point numbers that are represented by a first number of bits. One or more multiplexers are connected to the cache line(s). The multiplexer(s) is/are configured to select one of the floating-point numbers based on an integer index having a second number of bits that is less than the first number of bits. The integer index is a quantized representation of the selected one of the floating-point numbers. In some cases, a memory is configured to store a matrix of integer indices that are generated by quantizing floating-point numbers that represent data associated with a machine learning algorithm. The integer index is provided by the matrix of the integer indices.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
A processing system (200) and method (300) for executing arithmetic and conversion operations involving 16-bit brain floating-point (BF16)-formatted data are described. An instruction specifying either an arithmetic or conversion operation and a first data element in BF16 data format are received. For arithmetic operations, the exponents (130) of the data elements are aligned, and a result is generated using the aligned exponents. For conversion operations, the mantissa (140) of the BF16 data element is scaled based on its exponent, and the element is converted to a second data format, such as FP32 or a reduced-precision format, using precision-aware scaling and rounding. The result is stored in operations registers, such as for additional processing.
G06F 7/483 - Calculs avec des nombres représentés par une combinaison non linéaire de nombres codés, p. ex. nombres rationnels, système de numération logarithmique ou nombres à virgule flottante
G06F 7/501 - Semi-additionneurs ou additionneurs complets, c.-à-d. cellules élémentaires d'addition pour une position
Systems and techniques for providing co-issue of instructions utilize a scheduler (112) associated with a compute unit (124) to select one double-precision (i.e., 64-bit) instruction (201) and one single-precision (i.e., 32-bit) instruction (203) for issue to and execution at a compute unit. Each compute unit includes or is associated with one or more pairs of double-precision arithmetic logic units (ALUs) (210) and single-precision ALUs (208). The selected double-precision and single-precision instructions are associated with different threads or waves such that no dependency can exist between the two instructions. The selected single-precision instruction may perform address calculations or other memory tasks while the double-precision instruction may perform data computations that may be required for, e.g., matrix multiplication, or other machine learning functionality.
G06F 9/38 - Exécution simultanée d'instructions, p. ex. pipeline ou lecture en mémoire
G06F 9/48 - Lancement de programmes Commutation de programmes, p. ex. par interruption
G06F 7/57 - Unités arithmétiques et logiques [UAL], c.-à-d. dispositions ou dispositifs pour accomplir plusieurs des opérations couvertes par les groupes ou pour accomplir des opérations logiques
A hardware-accelerated siamese neural network device includes a first hardware-accelerated convolutional neural network (CNN) circuit configured to apply a certain weight to a first input at a specific moment of an operation. The hardware-accelerated siamese neural network device also includes a second hardware-accelerated CNN circuit configured to apply the certain weight to a second input at the specific moment of the operation. In addition, the hardware-accelerated siamese neural network device includes a classifier circuit configured to generate a score that represents a degree of similarity between the first input and the second input. Various other devices, systems, and methods are also disclosed.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
A processing system and method for executing arithmetic and conversion operations involving 16-bit brain floating-point (BF16)-formatted data are described. An instruction specifying either an arithmetic or conversion operation and a first data element in BF16 data format are received. For arithmetic operations, the exponents of the data elements are aligned, and a result is generated using the aligned exponents. For conversion operations, the mantissa of the BF16 data element is scaled based on its exponent, and the element is converted to a second data format, such as FP32 or a reduced-precision format, using precision-aware scaling and rounding. The result is stored in operations registers, such as for additional processing.
A processor includes an inference engine configured to dynamically adjust a first configuration associated with one or both of a processing system or an application at the processing system to a second configuration. The inference engine evaluates an impact of the second configuration on at least one performance metric and one or more states of the processing system. The inference engine further reverts to a previous stable configuration or maintains the second configuration based on the impact of the second configuration.
A processor protects a machine learning model (MLM) from unauthorized access. The processor employs a neural processing unit (NPU) to execute the MLM and implements decryption and encryption processes to decrypt the MLM and re-encrypt the MLM at different points along MLM storage and execution paths. Furthermore, the processor executes the encryption and decryption processes at different processing units and processing engines, thereby reducing the ability of malicious software to access the MLM. In addition, the processor protects buffers of the NPU from unauthorized access.
Methods, devices, and systems for information routing. A destination indication corresponding to data is received. The data is transmitted on a routing path based on routing information retrieved from a routing table bypass that includes a routing cache and a match mask. In some implementations, the destination indication comprises at least a portion of a network address. Some implementations include bypassing a routing table based on the routing information being present in the routing cache or the match mask. In some implementations, a total number of entries of the routing cache is fewer than a number of addresses in a working set. In some implementations, the match mask is dynamically programmable. In some implementations, an entry of the match mask includes a bit mask, a bit match and a routing destination associated with at least a portion of an address. In some implementations, the data comprises a packet.
Systems, methods, and non-transitory computer-readable media for dynamic voltage margin adjustment are disclosed. In some aspects, a system includes a processor, a power supply monitor (PSM) configured to measure voltage levels of the processor, and a voltage control module. The voltage control module may monitor minimum voltage levels reported by the PSM, compare the minimum voltage levels to one or more thresholds, and dynamically adjust a voltage margin applied to the processor based on the comparison. The voltage margin adjustment may be performed in real-time and may account for different workload characteristics, potentially optimizing processor performance and power efficiency.
Systems and techniques for providing mixed-precision matrix multiplication in multi-chiplet processors recognize different precision formats of matrices to be multiplied based on, e.g., parameters provided with instructions or start and end memory locations of the matrices. A plurality of different multiplication chains (208, 210, 212) are provided for different formats such that mixed-precision matrix multiplication can be performed using multiplication chains configured to handle multiplication of different precision formats. The multiplication chains are automatically selected based on the precision formats of the matrices to be multiplied, enabling programmers to utilize the chains without having to directly access the individual multiplication chains.
A flexible microscaling (MX) approach allows for the dynamic selection between multiple scale determination functions to compute a scaling factor that delivers lower deviation from an input data array (112, 302) when converting the input data array to an output data array (312) with reduced bit precision. The flexible MX approach includes selecting a first scale determination function from a plurality of scale determination functions (256-1, 256-2) based on a characteristic of an input data array having a first data format, computing a scaling factor (370) based on the first scale determination function, and converting the input data array into an output data array having a target data format based on the selected scaling factor. The output data array is then input to an artificial intelligence (AI) model for processing.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06F 7/483 - Calculs avec des nombres représentés par une combinaison non linéaire de nombres codés, p. ex. nombres rationnels, système de numération logarithmique ou nombres à virgule flottante
G06F 7/499 - Maniement de valeur ou d'exception, p. ex. arrondi ou dépassement
Systems and techniques for pin cutouts in memory modules (e.g., Dual In-line Memory Modules (DIMMs)) are described. A memory module includes one or more memory integrated circuits mounted on a printed circuit board. The memory module also includes at least one additional component, including, for example, a power manager integrated circuit, buffers, and inductors. The memory module also includes one or more rows of pins to provide an electrical connection to a motherboard. One or more rows include a pin cutout wherein some pins are removed or not placed in the row. The pin cutout allows a component to be mounted on the PCB in the foregone board area.
A technique for graphics processing is disclosed. The technique includes processing an allocation command for a graphics resource object, wherein the graphics resource object is a representation of a data store, the allocation command reserving address space for the data store as a sparse data store with uncommitted memory; and processing a commitment command for the graphics resource object that specifies a range of the data store and includes an indicator of commitment state, wherein processing the commitment command sets a commitment state of memory for the specified range according to the indicator.
Systems and methods are described for a workflow communications agent. The workflow communications agent receives a message that contains task-related content and parses the message to determine whether the task-related content includes predefined segments associated with one or more predefined instruction sets. If predefined segments are identified, the WCA initiates execution of the associated predefined instruction set(s). If no predefined segments are detected, the task-related content is provided to one or more large language models (LLMs) to determine whether the content can be associated with a predefined instruction set. In certain scenarios, the WCA generates new instruction sets based on contents of the task-related content if needed, and initiates execution of the predefined or newly generated instruction sets.
In an implementation, a computing device may include a processor having one or more cores. The computing device may also include a memory coupled to the processor, the memory having a page table. The computer device may further include where the processor is configured to receive a request to set a memory page to a fixed value, set, in a page table entry of the page table, a fixed page contents field to indicate the memory page contains the fixed value, and in response to a subsequent read request for the memory page, return the fixed value without accessing a physical page in the memory.
A memory controller for a memory channel includes first and second command sub-queues having corresponding first and second arbiters for selecting requests from respective ones of first and second queues for issuance as commands to the memory channel, and a power control circuit operable to place one of the first command sub-queue and the second command sub-queue into a low power mode while keeping another one of the first command sub-queue and the second command sub-queue active.
An apparatus and method for efficiently performing efficient data storage and data transfer of machine learning data. In various implementations, a host processing circuit of a computing system executes a machine learning (ML) application. The application includes a computational graph that indicates the computational order of the ML nodes, layers, and stages of the ML model. The host processing circuit translates function calls in the application to commands particular to an accelerator circuit. The accelerator circuit preloads weights to be used by ML nodes of the ML model by retrieving compressed weights from a storage device different from system memory. The accelerator circuit uses a streaming application programming interface (API) and bypasses the host processing circuit to retrieve the compressed weights from the storage device. The accelerator circuit decompresses the retrieved weights and executes the ML node using the decompressed weights.
One or more processors are configured to modify information representing popularities of entries stored in a first table in response to access requests to the entries. The popularities represent frequencies or likelihoods that corresponding ones of the entries were previously accessed. The processor is also configured to select a subset of the entries based on their popularities and copy the subset of the entries from the first table to a second table. In some cases, the entries in the first table stored vector embedding generated by a recommendation model. A filter, such as a Bloom filter, can be configured to selectively route data requests to the first table or the second table.
A feature referred to as “opacity micro-maps” provides an efficient representation of geometric detail for ray tracing primitives. One issue with opacity micro-maps is that in a highly parallel system with many rays referencing nearby opacity micro-map data in parallel, a great deal of unnecessary memory traffic may be generated. To combat this unnecessary memory traffic, a mechanism is provided herein for reducing the number of redundant memory requests. According to this mechanism, opacity micro-map circuitry maintains an indication of pending memory requests for opacity micro-map data. If a new request for opacity micro-map evaluation occurs and the data required for that new request is at the same address as the pending memory request, then no additional memory request is generated for the new request. When the data from the pending memory request is returned from memory, that data is used to satisfy all evaluation requests that are outstanding.
Data for nodes of a bounding volume hierarchy can be represented in fixed-point format for compression. For good precision, fixed-point bounds are stored along with each node, where these bounds represent the minimum and maximum representable numbers in the format. While this provides good compression, there are ways in which this technique can produce undesirable results. For example, if the size of many or most of the bounding volumes in the fixed-point number space is significantly smaller than the bounds, then the representation of such bounding volume in the fixed-point space may be unnecessarily large, resulting in a large number of false positive intersections during BVH traversal. Thus, techniques are provided herein for setting certain bounding volumes as “always hit” in order to eliminate the bounding volumes from inclusion in the fixed-point bounds. This shrinks the fixed-point bounds, giving better precision for smaller bounding volumes.
Systems and techniques for low-power prefetching are described. In one example, a processor includes a cache system having a hierarchy of one or more cache levels and a prefetcher associated with a cache level of the cache system. The prefetcher determines a memory level from which data or instructions associated with a prefetch request is retrieved. In response to the data or instructions being retrieved from a level three cache, a lower cache level, or system memory, the prefetcher stores the data or instructions associated with the prefetch request in a level two cache. The described techniques reduce cache pollution in the level one cache without introducing additional storage overhead.
G06F 12/0862 - Adressage d’un niveau de mémoire dans lequel l’accès aux données ou aux blocs de données désirés nécessite des moyens d’adressage associatif, p. ex. mémoires cache avec pré-lecture
G06F 12/0811 - Systèmes de mémoire cache multi-utilisateurs, multiprocesseurs ou multitraitement avec hiérarchies de mémoires cache multi-niveaux
G06F 12/0875 - Adressage d’un niveau de mémoire dans lequel l’accès aux données ou aux blocs de données désirés nécessite des moyens d’adressage associatif, p. ex. mémoires cache avec mémoire cache dédiée, p. ex. instruction ou pile
The disclosed device provides smart retimer features for a bus that is compatible with slower speed devices, such as devices using an older generation protocol for the bus. The smart retimer device can convert data packets into different formats and utilize available lanes for more efficient use of available bandwidth. Various other methods, systems, and computer-readable media are also disclosed.
Systems and methods described for choosing optimal resizing images in an image dataset are described. An image or a region of interest (ROI) of the image is chosen and complexity of the image or ROI is computed using one or more methods. Different resizing methods are used for less complex images or ROI, while images with more complex image or ROI are resized differently. The complexity is precomputed for different tiles of images and embedded as metadata within the images. In some cases, the complexity estimates can be calculated dynamically, e.g., when no preprocessed image data is available.
Dynamically modifying voltage regulator modes based on a computing system workload is described. System management circuitry assesses activity and energy demands of a computing system’s hardware components for a given workload and adapts voltage regulator configuration settings to ensure that the voltage regulator is operating in a mode that is optimized for the workload. In some implementations, modifying voltage regulator configuration settings causes a voltage regulator to transition from a current mode to a different mode. Alternatively or additionally, modifying configuration settings causes the voltage regulator to continue operating in a current mode with a different rate of voltage mitigation.
Selective data compression for non-critical memory requests is described. In accordance with the described techniques, memory request packets are generated and communicated for retrieval of data from computer memory, and data packets including the retrieved data are communicated to a hardware compression engine. The compression engine selectively compresses the retrieved data and the compressed data is communicated through an interconnect architecture for decompression by a decompression engine. The compression of the data decreases communication congestion within the interconnect architecture. The memory request packets are configurable to include metadata providing hints to the compression engine to guide selective compression of data.
Techniques are disclosed for transposing and loading 6-bit floating-point matrix data into operations registers of a processing unit. A matrix comprising 6-bit floating-point data elements is received from memory, in which the matrix is stored in either a row-major or column-major layout. A matrix transposition loading operation is performed during the loading process, rearranging the matrix elements into a transposed layout (e.g., converting column-major to row-major). The transposed matrix is stored in operations registers for use in parallel processing tasks. The process may include caching partially stored data elements from memory and combining them with subsequently retrieved data to complete the transposition.
Systems and techniques for providing mixed-precision matrix multiplication in multi-chiplet processors recognize different precision formats of matrices to be multiplied based on, e.g., parameters provided with instructions or start and end memory locations of the matrices. A plurality of different multiplication chains are provided for different formats such that mixed-precision matrix multiplication can be performed using multiplication chains configured to handle multiplication of different precision formats. The multiplication chains are automatically selected based on the precision formats of the matrices to be multiplied, enabling programmers to utilize the chains without having to directly access the individual multiplication chains.
Systems and techniques for providing co-issue of instructions utilize a scheduler associated with a compute unit to select one double-precision (i.e., 64-bit) instruction and one single-precision (i.e., 32-bit) instruction for issue to and execution at a compute unit. Each compute unit includes or is associated with one or more pairs of double-precision arithmetic logic units (ALUs) and single-precision ALUs. The selected double-precision and single-precision instructions are associated with different threads or waves such that no dependency can exist between the two instructions. The selected single-precision instruction may perform address calculations or other memory tasks while the double-precision instruction may perform data computations that may be required for, e.g., matrix multiplication, or other machine learning functionality.
Metadata storage for a stacked die configuration is described. In accordance with the described techniques, a system includes a first die that has one or more processor cores and shared cache. The system also includes a second die with metadata storage. In response to a context switch or power state transition for the one or more processor cores, a metadata controller causes metadata associated with the one or more processor cores to be saved in the metadata storage on the second die. The metadata controller also loads metadata associated with the new context (e.g., a different process or thread) or the new power state from the metadata storage.
G06F 12/0862 - Adressage d’un niveau de mémoire dans lequel l’accès aux données ou aux blocs de données désirés nécessite des moyens d’adressage associatif, p. ex. mémoires cache avec pré-lecture
G06F 9/46 - Dispositions pour la multiprogrammation
G06F 12/0811 - Systèmes de mémoire cache multi-utilisateurs, multiprocesseurs ou multitraitement avec hiérarchies de mémoires cache multi-niveaux
89.
Bandwidth Management for Real-Time and Best-Effort Clients Under Loaded System Conditions
A power manager of an apparatus exposes an application programming interface (API) usable for applications to specify priority and quality-of-service (QoS) parameters (e.g., bandwidth requirements) for a workload. An application, for instance, specifies the priority and QoS parameters for a workload to be processed using a hardware compute unit. The power manager employs the priority and QoS parameters to configure the bandwidth allocation to access a memory system. In particular, the bandwidth allocation and prioritization are dynamically extended to real-time and best-effort workloads to satisfy specified QoS parameters for inference workloads and improve user experiences.
An accelerator unit (AU) includes a compute unit configured to execute a first workgroup of a first kernel and a set of compute unit resources allocated to the first workgroup. Concurrently with the compute unit executing the first workgroup, a scheduling circuitry of the AU receives a resource termination hint indicating that the first workgroup is going to end the use of portions of compute unit resources allocated to the first workgroup. In response to this resource termination hint, the scheduling circuitry provisionally allocates these portions of compute unit resources to a second workgroup of a second kernel and begins execution of a portion of the second workgroup. After the first workgroup releases the portions of compute unit resources, the scheduling circuitry fully allocates the portions of compute unit resources to the second workgroup and executes a remainer of the second workgroup.
An important part of graphics rendering is determining lighting, which affects the illumination of objects in a scene. A lighting calculation includes several components, including direct lighting, which is light reflected from a light source across a surface into a camera. Calculating direct lighting is a processing intensive operation since, among other things, it must be determined whether each light source is visible at a particular surface. A technique is provided herein whereby a trained neural network is used to assist with direct lighting calculations. The trained neural network is trained using data from the current scene being rendering. The network accepts a point within the scene as a query and outputs lighting results for the point within the scene. In various examples, the lighting results indicate whether a light is visible from that point or the strength of the light as viewed from that point.
A processing system executing a generative artificial intelligence model generates keys and values (KV vectors) just in time for consumption by a layer of the model by selectively recomputing keys and values that were computed in a previous layer of the model rather than storing the keys and values for consumption by subsequent layers.
An apparatus and method for efficiently scheduling instructions for a parallel data processing circuit. In various implementations, a computing system includes a variety of types of processing circuits with two or more capable of executing a same type of task. A hardware component, such as a processing circuit, accesses a command buffer. The processing circuit reads, in the command buffer, a predicate command corresponding to the next task to execute. The processing circuit checks the predicate memory location corresponding to the next task to verify whether another hardware component has started the next task. If any other hardware component has begun executing the task, then the processing circuit discards the task from its command buffer. Otherwise, if no other hardware component has begun executing the task, then the processing circuit updates the predicate memory location corresponding to the task to specify the task has begun execution.
A technique for generating backward optical flow data includes identifying a forward velocity value of a picture element at first coordinates in a subsequent frame buffer that stores forward velocity values for a subsequent frame. The technique identifies second coordinates in a previous frame buffer that stores backward velocity values of a previous frame by adjusting the first coordinates based on a negated version of the forward velocity value. A backward velocity value for the picture element at the second coordinates in the previous frame buffer is set to the negated version of the forward velocity value of the picture element at the first coordinates in the subsequent frame buffer. An image is rendered in accordance with the backward velocity value for the picture element at the second coordinates in the previous frame buffer.
Embodiments herein identify anchor points in an isometric image (e.g., bottom, left, right, and top anchor points) which can be used to adjust the image so it has a desired perspective (e.g., where the angles between the x, y, and z axes are the same). To do so, a system determines horizontal adjustments for the anchor points using polynomial interpolation. The system then determines vertical adjustments for the anchor points using cubic polynomial interpolation. Doing so adjusts or stretches the image so that it has the desired orientation. In this manner, isometric images (e.g., AI-generated images) that each have a different orientation can be separately adjusted to have the same orientation so they, for example, each can be placed in grids with the same dimensions for an isometric game.
A chip package for a computer system includes a substrate including a glass substrate core, first redistribution layers disposed on a first surface of the glass substrate core, and second redistribution layers disposed on a second surface of the glass substrate core. The chip package further includes an interposer including a glass interposer core, third redistribution layers disposed on a first surface of the glass interposer core, and fourth redistribution layers disposed on a second surface of the glass interposer core. Further, the chip package includes a first chip die mounted to a first surface of the interposer.
Accelerating and improved fairness for semaphores is described. In one or more implementations, a computing device includes a plurality of cores and one or more computing resources, exclusive access to which is limited, e.g., to a single core at a time. In one or more implementations, the computing device also includes core selection circuitry configured to identify a first core from the plurality of cores exclusively accessing the one or more limited computing resources and configured to instruct a second core from the plurality of cores to release the exclusive access of the first core to the one or more computing resources.
A feature referred to as "opacity micro-maps" provides an efficient representation of geometric detail for ray tracing primitives. One issue with opacity micro-maps is that in a highly parallel system with many rays referencing nearby opacity micro-map data in parallel, a great deal of unnecessary memory traffic may be generated. To combat this unnecessary memory traffic, a mechanism is provided herein for reducing the number of redundant memory requests. According to this mechanism, opacity micro-map circuitry maintains an indication of pending memory requests for opacity micro-map data. If a new request for opacity micro-map evaluation occurs and the data required for that new request is at the same address as the pending memory request, then no additional memory request is generated for the new request. When the data from the pending memory request is returned from memory, that data is used to satisfy all evaluation requests that are outstanding.
The disclosed circuit can select a key index in response to a memory request including a physical address. The physical address points to a location of a graphics processing unit memory that is encrypted. The circuit can forward the selected key index with the physical address to a memory controller of the graphics processing unit memory. The memory controller can complete the memory request using a key associated with the key index. Various other methods, systems, and computer-readable media are also disclosed.
G06F 12/14 - Protection contre l'utilisation non autorisée de mémoire
G06F 21/79 - Protection de composants spécifiques internes ou périphériques, où la protection d'un composant mène à la protection de tout le calculateur pour assurer la sécurité du stockage de données dans les supports de stockage à semi-conducteurs, p. ex. les mémoires adressables directement
The disclosed device includes a data fabric controller that can identify which computing nodes are more biased to remote memory traffic and which computing nodes are more biased to local memory traffic. In response to detecting a load condition on a remote memory device, the data fabric controller can throttle one of the computing nodes that are biased to remote memory traffic. Various other methods, systems, and computer-readable media are also disclosed.