Embodiments herein describe a system including a direct memory access (DMA) engine configured to receive first data stored in a local memory and an artificial intelligence (AI) engine configured to receive second data bypassing the DMA engine. The first data includes matrix-matrix instructions and the second data includes vector-matrix instructions. The second data is configured to be directly sent to the AI engine to avoid additional overhead of DMA data transfer latency, reduce memory bandwidth usage, and minimize power consumption.
G06F 13/28 - Gestion de demandes d'interconnexion ou de transfert pour l'accès au bus d'entrée/sortie utilisant le transfert par rafale, p. ex. acces direct à la mémoire, vol de cycle
G06F 9/52 - Synchronisation de programmesExclusion mutuelle, p. ex. au moyen de sémaphores
An adaptive voltage assist circuit includes a voltage comparison circuit with a first input configured to receive a supply voltage, and a second input configured to receive a threshold voltage. The supply voltage may be configured to be changed dynamically. The voltage comparison circuit is configured to generate at least one output signal indicating whether the supply voltage is above the threshold voltage. An assist logic circuit is configured to use the at least one output signal to generate a control signal indicating a voltage assist configuration for a memory circuit, such as read assist and/or write assist configurations. The memory circuit may be a volatile memory circuit, such as a static random-access memory (SRAM) circuit. The adaptive voltage circuit may be included in an integrated circuit that includes the memory circuit and additional circuits that operate using the supply voltage.
Embodiments herein describe an artificial intelligence (AI) engine including first functional circuitry configured to perform a first set of instructions that include a matrix operation and second functional circuitry configured to perform a second set of instructions that include a vector operation, where the second set of instructions are concurrently performed with the first set of instructions. The first set of instructions include instructions for an operation to perform based on two or more matrices and the second set of instructions include instructions for an operation to perform based on one or more vectors. The first set of instructions include instructions for an operation to perform based on a matrix and a vector and the second set of instructions include instructions for an operation to perform based on one or more vectors.
Embodiments herein describe an artificial intelligence (AI) engine including a register file and a multiplier configured to support a superset of a data type. The register file stores multiple variants of the data type. The AI engine includes logic to expand multiplication operations performed by the multiplier. The data type is a floating-point data type, where the superset of the data type includes a maximum value for a mantissa and a maximum value for an exponent.
G06F 5/01 - Procédés ou dispositions pour la conversion de données, sans modification de l'ordre ou du contenu des données maniées pour le décalage, p. ex. la justification, le changement d'échelle, la normalisation
Embodiments herein describe a processor system that includes an integrated, adaptive accelerator. In one embodiment, the processor system includes multiple core complex chiplets that each contain one or processing cores for a host CPU. In addition the processor system includes an accelerator chiplet. The processor system can assign one or more of the core complex chiplets to the accelerator chiplet to form an IO device while the remaining core complex chiplets form the CPU for the host. In this manner, rather than the accelerator and the CPU having independent computer resources, the accelerator can be integrated into the processor system of the host so that hardware resources can be divided between the CPU and the accelerator depending on the needs of the particular application(s) executed by the host.
A chip package and method for fabricating the same are provided that includes embedded off-die inductors coupled in series. One of the off-die inductors is disposed in a redistribution layer formed on a bottom surface of an integrated circuit (IC) die. The other of the series connected off-die inductors is disposed in a substrate of the chip package. The substrate may be either an interposer or a package substrate.
A processor includes at least one processing elements. The at least one processing element is configured to generate a first optimized input matrix and a second optimized input matrix. The first optimized input matrix is generated based on copying elements of a column or row that includes a target value in a first input matrix to an unused column or row in the first input matrix, and scaling the elements prior or subsequent to copying elements. The second optimized input matrix is generated based on based on copying elements of a corresponding row or column in a second input matrix to an unused row or column in the second input matrix. The at least one processing element is further configured to generate an output matrix based on multiplying the first optimized input matrix and the second optimized input matrix, and perform one or more actions based on the output matrix.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
Disclosed devices, systems, and methods may enhance communication protocols for low latency applications. Systems may include a device and a host interconnected by a high-performance interconnect and/or communication link. The device may comprise a transmitter, a receiver, and a control unit that may manage a credit-based flow control mechanism. In some aspects, the device may initiate a push write request, send a data header with an identifier (UQID) matching the push write request identifier (CQID), and transmit the data payload. The host may receive the push write request, match the UQID with the CQID, perform the write operation, and send a completion message back to the device. The method may involve ensuring sufficient credits before initiating the push write transaction, which may help prevent data loss and ensure reliable delivery. The push write mechanism may reduce the number of link traversals required for device-to-host memory writes, potentially lowering overall latency.
A device may include an array of data processing engines (DPEs) on a die and an event broadcast network. Each of the DPEs includes a core, a memory module, event logic in at least one of the core or the memory module, and an event broadcast circuitry coupled to the event logic. The event logic is capable of detecting an occurrence of one or more events in the core or the memory module. The event broadcast circuitry is capable of receiving an indication of a detected event detected by the event logic. The event broadcast network includes interconnections between the event broadcast circuitry of the DPEs. Detected events can trigger or initiate various responses, such as debugging, tracing, and profiling.
A processor includes at least one processing elements. The at least one processing element is configured to generate a first optimized input matrix and a second optimized input matrix. The first optimized input matrix is generated based on copying elements of a column or row that includes a target value in a first input matrix to an unused column or row in the first input matrix, and scaling the elements prior or subsequent to copying elements. The second optimized input matrix is generated based on based on copying elements of a corresponding row or column in a second input matrix to an unused row or column in the second input matrix. The at least one processing element is further configured to generate an output matrix based on multiplying the first optimized input matrix and the second optimized input matrix, and perform one or more actions based on the output matrix.
G06F 5/01 - Procédés ou dispositions pour la conversion de données, sans modification de l'ordre ou du contenu des données maniées pour le décalage, p. ex. la justification, le changement d'échelle, la normalisation
11.
MODEL LEVEL DEBUGGING OF MACHINE LEARNING DESIGNS ON NEURAL PROCESSING UNITS
Model level debugging of a machine learning design includes compiling the machine learning design for execution on target hardware using a compiler. Metadata for the machine is generated. The metadata specifies a mapping of buffers of the machine learning design to a plurality of memory levels of a memory architecture of the target hardware correlated with boundaries of the machine learning design. While running the machine learning design, debug data is dumped from the plurality of memory levels of the memory architecture based on the boundaries. The debug data is correlated with the boundaries of the machine learning design based on the metadata.
Devices, systems, and methods manage activity in die-to-die links. A link controller places a die-to-die link in a partially active state, with some lane groups active and others idle, adapting to changing conditions without full link retraining. Lane groups can independently enter an electrical idle state, reducing power consumption without interrupting data transmission over active lanes. Idle lanes can be retrained and reactivated without affecting active lanes. Control messages, may manage lane group states, ensuring synchronized transitions between active and idle states. This approach optimizes power management and maintains data transmission efficiency in die-to-die interconnects.
Examples herein describe techniques for reducing the amount of memory used during weight sparsity. When decompressing the weights, the uncompressed weight data typically has many zero values. By knowing the location of these zero values (e.g., their indices in a weight matrix), the processor core can prune some of the activations (e.g., logically reduce the size of the activation matrix) which improves the efficiency of the processor core. In embodiments herein, the processor core includes logic for identifying the indices of the non-zero value after decompressing the compressed weights. These indices can then be used to prune the activations to improve the efficiency of the processor core.
Disclosed devices, systems, and methods may enhance communication protocols for low latency applications. Systems may include a device and a host interconnected by a high-performance interconnect and/or communication link. The device may comprise a transmitter, a receiver, and a control unit that may manage a credit-based flow control mechanism. In some aspects, the device may initiate a push write request, send a data header with an identifier (UQID) matching the push write request identifier (CQID), and transmit the data payload. The host may receive the push write request, match the UQID with the CQID, perform the write operation, and send a completion message back to the device. The method may involve ensuring sufficient credits before initiating the push write transaction, which may help prevent data loss and ensure reliable delivery. The push write mechanism may reduce the number of link traversals required for device-to-host memory writes, potentially lowering overall latency.
A flexible microscaling (MX) approach allows for the dynamic selection between multiple scale determination functions to compute a scaling factor that delivers lower deviation from an input data array (112, 302) when converting the input data array to an output data array (312) with reduced bit precision. The flexible MX approach includes selecting a first scale determination function from a plurality of scale determination functions (256-1, 256-2) based on a characteristic of an input data array having a first data format, computing a scaling factor (370) based on the first scale determination function, and converting the input data array into an output data array having a target data format based on the selected scaling factor. The output data array is then input to an artificial intelligence (AI) model for processing.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
G06F 7/483 - Calculs avec des nombres représentés par une combinaison non linéaire de nombres codés, p. ex. nombres rationnels, système de numération logarithmique ou nombres à virgule flottante
G06F 7/499 - Maniement de valeur ou d'exception, p. ex. arrondi ou dépassement
A processing system is configured to generate a three-dimensional (3D) Gaussian map representing at least a portion of an environment surrounding the processing system. For example, a capture device of the processing device first captures a set of frames each including color data. The processing system then implements a visual odometry (VO) tracking model which samples patches from this set of frames so as to generate a point cloud and pose data representing the location and orientation of the patches within the environment. Further, the processing system implements a Gaussian Mapping model that generates a set of Gaussians from the point cloud which the processing system then uses to populate a 3D Gaussian map.
High-level synthesis generation of multiplexer logic includes generating, using computer hardware, an intermediate representation of a design for an integrated circuit. The design is specified in high-level programming language source code. Selected logic within the intermediate representation is detected by the computer hardware and converted into a sparsemux node. A selected core is chosen from a plurality of cores to implement the sparsemux node. The computer hardware is capable of choosing the selected core based on a label encoding format and a default input of the sparsemux node. A circuit design is generated by the computer hardware. One or more operations of the selected core are scheduled based on a timing model corresponding to the selected core.
G06F 30/327 - Synthèse logiqueSynthèse de comportement, p. ex. logique de correspondance, langage de description de matériel [HDL] à liste d’interconnections [Netlist], langage de haut niveau à langage de transfert entre registres [RTL] ou liste d’interconnections [Netlist]
A flexible microscaling (MX) approach allows for the dynamic selection between multiple scale determination functions to compute a scaling factor that delivers lower deviation from an input data array when converting the input data array to an output data array with reduced bit precision. The flexible MX approach includes selecting a first scale determination function from a plurality of scale determination functions based on a characteristic of an input data array having a first data format, computing a scaling factor based on the first scale determination function, and converting the input data array into an output data array having a target data format based on the selected scaling factor. The output data array is then input to an artificial intelligence (AI) model for processing.
Embodiments herein relate to applying a Boolean matrix factorization operation to selected neurons, represented as truth tables, in LUT based neural networks. Applying the Boolean matrix factorization operation replaces the original truth table representing the neuron of the neural network with a simplified, approximated truth table derived from the original truth table. The neural network is then simulated, and the cost of the simulated result, as evaluated by a cost function which may for instance incorporate accuracy and resource cost, is compared to a cost threshold value to decide whether to replace the original neuron with its approximated representation.
Alignment detection circuitry is disclosed that include a buffer configured to output a data stream of multiplexed groups of symbols from multiple data lanes and a set of correlators configured to determine physical link skew from one data lane of the multiple data lanes, wherein the physical link skew from the one data lane is known to equal the physical link skew on each data lane, assuming each data lane is transmitted over the same physical link. The alignment detection circuitry may also include anti¬ aliasing circuitry configured to extract a unique marker of the alignment marker to detect common marker aliasing by comparing the extracted unique marker to a known unique marker. Anti-aliasing circuitry may also be configured to use common marker detection status of adjacent data lanes of the multiple data lanes to detect common marker aliasing determined by a detection delay between the adjacent data lanes.
Calibration of wavelength division multiplexing (WDM) optical micro-ring modulators (MRMs) may include providing multiple optical signals to a waveguide having multiple MRMs, sweeping heater settings of each MRMs, from high to low, recording drop-port outputs of the first MRM for the corresponding heater settings, while other MRMs are rendered transparent, determining a percentage of the peak-drop port outputs and corresponding off-peak heater settings, discarding off-peak heater settings that are below a tracking margin threshold, and selecting calibrated heater settings for the MRMs from remaining ones of the off-peak heater settings of the respective MRMs. The calibrated heater settings may be selected based on a greedy placement method. Alternatively, additional data may be generated by determining remaining heater settings that cause collisions between adjacent pairs of MRMs, and selecting the calibrated heater settings based on an ordered placement or a sorted placement of the collision settings.
G02F 1/01 - Dispositifs ou dispositions pour la commande de l'intensité, de la couleur, de la phase, de la polarisation ou de la direction de la lumière arrivant d'une source lumineuse indépendante, p. ex. commutation, ouverture de porte ou modulationOptique non linéaire pour la commande de l'intensité, de la phase, de la polarisation ou de la couleur
G02F 1/025 - Dispositifs ou dispositions pour la commande de l'intensité, de la couleur, de la phase, de la polarisation ou de la direction de la lumière arrivant d'une source lumineuse indépendante, p. ex. commutation, ouverture de porte ou modulationOptique non linéaire pour la commande de l'intensité, de la phase, de la polarisation ou de la couleur basés sur des éléments à semi-conducteurs ayant des barrières de potentiel, p. ex. une jonction PN ou PIN dans une structure de guide d'ondes optique
22.
TRANSFORM BLOCK FLOATING POINT VALUES TO OTHER BLOCK FLOATING POINT FORMATS
Embodiments herein describe an artificial intelligence (AI) engine including a register file including a block floating point configured to be converted to a block size that aligns with hardware capabilities by dividing the block floating point into multiple sub-blocks and replicating an exponent of the block floating point for each of the multiple sub-blocks. In one example, the block floating point includes 32 elements and the multiple sub-blocks are 4 sub-blocks with 8 elements each. In another example, the block floating point includes 32 elements and the multiple sub-blocks are 2 sub-blocks with 16 elements each.
G06F 7/483 - Calculs avec des nombres représentés par une combinaison non linéaire de nombres codés, p. ex. nombres rationnels, système de numération logarithmique ou nombres à virgule flottante
23.
ADAPTIVE QUANTIZATION AND COMPRESSION FOR VARIABLE-LENGTH COLLECTIVE COMMUNICATION
A processing system dynamically and selectively quantizes or compresses variable-length input data for a collective operation to fit within a predetermined limit, executes the collective operation on the compressed data, and converts the data back to variable length results by dequantizing or decompressing the results of the collective operation.
Embodiments herein describe a sparse-dense-sparse (SDS) process that achieves a better pruning scheme that benefits from pruning-friendliness relative to one-shot pruning schemes. The SDS process performs a first pruning to generate a sparse ML model followed by reconstruction to generate a re-dense ML model, followed by a second pruning to generate another sparse ML model. By pruning a ML model and then re-constructing the ML model, the ML model can be made more pruning-friendly by performing data and/or weight regularization. As a result, performing the second pruning in the SDS process can result in a smoother weight distribution and lower perplexity relative to one-shot pruning.
Embodiments herein describe relaxation oscillator circuits for physical unclonable function (PUF) circuits and/or for providing clocks to multiple dies. The relaxation oscillator circuit generates a clock having a frequency that is based in part on random process variations of the passive components of the relaxation oscillator circuit. The relaxation oscillator circuit may be distributed amongst multiple dies of an integrated circuit device, to incorporate additional sources of randomness/entropy, enhance security, to provide clocks to the multiple dies. The relaxation oscillator circuits include differential and single-ended relaxation oscillator circuits. Resistive and/or capacitive components may be configurable, which may be useful for testing purposes, altering PUF codes, and/or generating time-multiplexed clocks.
H03B 5/24 - Élément déterminant la fréquence comportant résistance, et soit capacité, soit inductance, p. ex. oscillateur à glissement de phase l'élément actif de l'amplificateur étant un dispositif à semi-conducteurs
H03K 3/3565 - Circuits bistables bistables avec hystérésis, p. ex. déclencheur de Schmitt
H04L 9/32 - Dispositions pour les communications secrètes ou protégéesProtocoles réseaux de sécurité comprenant des moyens pour vérifier l'identité ou l'autorisation d'un utilisateur du système
26.
Stress-reduced package substrate and method of forming the same
Disclosed herein is a package substrate and a method for fabricating the same. In one example, a package substrate includes a core having an outer edge and a first plurality of first interconnect layers disposed on the core. The first plurality of interconnect layers disposed on the core include an outermost dielectric layer disposed farthest from the core. The outermost dielectric layer has an edge that is recessed from the edge of the core.
Embodiments herein describe a system including a plurality of network resources providing communication between a plurality of source end points and a plurality of destination end points, where a first source end point transmits data to a first destination end point by using information about a size, route, and timing of data exchanges to pre-allocate resources in the system for a predefined time frame and use a common time reference among the first end points and the second end points to commence transmission of the data at a time when an allocation time frame begins.
Embodiments herein describe a hardware accelerator that includes multiple power or clock domains. For example, the hardware accelerator can include an array of data processing engines (DPEs) where different subsets of the DPEs (e.g., different columns, rows, or blocks) are disposed in different power or clock domains within the hardware accelerator. When one or more subsets of the DPEs are idle (e.g., the hardware accelerator has not assigned any tasks to those DPEs), the accelerator can deactivate the corresponding power or clock domain (or domains), which deactivates the DPEs in those domains while the DPEs in the other power or clock domains remain operational. As such, idle DPEs can be deactivated to conserve energy while DPEs with work can remain operational.
G06F 1/3237 - Économie d’énergie caractérisée par l'action entreprise par désactivation de la génération ou de la distribution du signal d’horloge
G06F 1/3228 - Surveillance d’exécution de tâches, p. ex. par utilisation de temporisations d’attente, de commandes d’arrêt ou de commandes d’attente
G06F 1/3287 - Économie d’énergie caractérisée par l'action entreprise par la mise hors tension d’une unité fonctionnelle individuelle dans un ordinateur
A modular and scalable high-performance compression system includes multiple compression units (CUs), each including a compression circuit, an input buffer, and a history buffer, and further including input circuitry that loads the input buffers and the history buffers with sub-blocks of data blocks. The compression circuits identify segments of the sub-blocks of the corresponding input buffers that match segments of the sub-blocks of the corresponding history buffers, and encode the identified segments based on positions/offsets of the matching segments within the history buffers. The system may include multiple selectable compression modes, and may compress multiple data blocks in parallel, based on the same or differing compression modes. The compression modes may include a high throughput mode and a high compression ratio mode, and may further include dynamic selection mode that selects a compression mode based on a compression ratio, available resources, and/or other criteria.
A multi-die ring oscillator with split entropy for security and physical unclonable function (PUF) applications includes a PUF circuit that includes a current source circuit that sources currents from a gated supply voltage and controls the currents based on an analog voltage, a voltage controller (e.g., a differential-input OTA) that controls the analog voltage based on first and second ones of the currents, and a ring oscillator that outputs a clock based on a third one of the currents, where a frequency of the clock is based on the third current and random variations of components of the PUF circuit. The voltage controller and the ring oscillator may be placed on a first die, and the current source circuit may be placed on a second die. The PUF circuit may be configurable to alter the clock frequency and/or to provide the clock with time-multiplexed frequencies.
G06F 21/73 - Protection de composants spécifiques internes ou périphériques, où la protection d'un composant mène à la protection de tout le calculateur pour assurer la sécurité du calcul ou du traitement de l’information par création ou détermination de l’identification de la machine, p. ex. numéros de série
G06F 21/72 - Protection de composants spécifiques internes ou périphériques, où la protection d'un composant mène à la protection de tout le calculateur pour assurer la sécurité du calcul ou du traitement de l’information dans les circuits de cryptographie
H04L 9/32 - Dispositions pour les communications secrètes ou protégéesProtocoles réseaux de sécurité comprenant des moyens pour vérifier l'identité ou l'autorisation d'un utilisateur du système
31.
LANE ALIGNMENT FOR SYMBOL MULTIPLEXED WIRED COMMUNICATION SYSTEMS
Alignment detection circuitry is disclosed that include a buffer configured to output a data stream of multiplexed groups of symbols from multiple data lanes and a set of correlators configured to determine physical link skew from one data lane of the multiple data lanes, wherein the physical link skew from the one data lane is known to equal the physical link skew on each data lane, assuming each data lane is transmitted over the same physical link. The alignment detection circuitry may also include anti-aliasing circuitry configured to extract a unique marker of the alignment marker to detect common marker aliasing by comparing the extracted unique marker to a known unique marker. Anti-aliasing circuitry may also be configured to use common marker detection status of adjacent data lanes of the multiple data lanes to detect common marker aliasing determined by a detection delay between the adjacent data lanes.
Data spray techniques are used to implement fault resistant data detection using a plurality of data detectors. Each of the plurality of data detectors is coupled, at sequential times, to a lane of data bytes during transfers thereof from a data source to a data destination. Each data detector is individually associated with the lane of data bytes at sequential time slots representing each data byte transfer. All bits of the bytes being transferred on the lane are examined individually by the data detectors in determining if the data byte has been programmed for a security key code by detecting a certain logic state in at least one bit thereof. Associating each of the plurality of data detectors during data byte transfers improves security key detection by reducing the probability of a single defective data detector leading to an erroneous conclusion of the status of the data.
G06F 21/79 - Protection de composants spécifiques internes ou périphériques, où la protection d'un composant mène à la protection de tout le calculateur pour assurer la sécurité du stockage de données dans les supports de stockage à semi-conducteurs, p. ex. les mémoires adressables directement
33.
DISTRIBUTED ON-CHIP KEY-VALUE STORE CIRCUIT ARCHITECTURE
An electronic system includes a key memory circuit capable of storing a plurality of keys and outputting an address for a key of the plurality of keys matched to a search key. The electronic system includes a memory manager circuit capable of tracking free memory addresses in the key memory circuit. The electronic system includes a state tracker circuit including a value memory circuit. The value memory circuit is capable of storing values of state information and is separate from the key memory circuit. The electronic system includes an insert-delete circuit capable of passing the address from the key memory circuit to the state tracker circuit, performing insert and delete operations on the key memory circuit, and providing free memory addresses of the key memory circuit to the memory manager circuit.
A method includes, for speculative decoding using a draft model that generates look ahead tokens for a target model, splitting an input dataset into a first set and a second set. The method also includes, using the first set, computing an expected number of look ahead tokens accepted by the target model per call based on a first value indicative of a quantity of look ahead tokens generated by the draft model per call to the target model, and modifying the quantity of look ahead tokens generated by the draft model to a second value based on the expected number of look ahead tokens accepted by the target model. Then, the method includes, with the second set, performing speculative decoding with a look ahead window size based on the second value.
Embodiments herein describe a system including a plurality of network resources providing communication between a plurality of source end points and a plurality of destination end points, where a first source end point transmits data to a first destination end point by using information about a size, route, and timing of data exchanges to pre-allocate resources in the system for a predefined time frame and use a common time reference among the first end points and the second end points to commence transmission of the data at a time when an allocation time frame begins.
Embodiments herein describe a system including a plurality of network resources providing data transmission between a first end point and a second end point by making a request for a network resource reservation to reserve resources on an end-to-end path during a time frame for the data transmission, selecting a routing table, populating the routing table with a limited number of entries, activating the data transmission at a start time of the time frame to an input port corresponding to the routing table, in response to activating the data transmission, forwarding data based on the entries, and determining an end time of the time frame to deactivate the data transmission from the first end point.
Machine learning (ML) tools for analyzing semiconductor yield analysis data, such as to infer information related to defects based on feature values of the yield diagnostic data, where the yield diagnostic data includes wafer-level test data, wafer-region based defect densities (DDs), and functional circuit block-based DDs. The wafer region based DDs are a function of numbers of defects within regions of the wafers and physical areas of the regions. The functional circuit block based DDs are a function of numbers of defects of functional circuit blocks and physical areas of the functional circuit blocks. The ML tools may include a supervised learning-based classification model, such as a gradient boosting classifier, an unsupervised outlier detection model, such as a hierarchical density-based spatial clustering of applications with noise model, an unsupervised clustering model, and/or an unsupervised associative model.
The embodiments herein describe techniques for performing ML compilation using a unified interface that combines different processors in a heterogeneous processing system which allows for intelligent partitioning of a ML model. Unlike prior solutions which rely on user preferences to assign the ML model, the unified interface can violate or break the user preferences when partitioning the ML model. The unified interface can receive information from the processors (e.g., a NPU, CPU, GPU, etc.) and determine the capabilities, current workload, power metrics, subgraphs of the ML model they can execute, and the like. With this information, the unified interface can intelligently choose when to violate or break the user-entered priority based instructions.
The embodiments herein describe techniques for performing ML compilation using a unified interface that combines different processors in a heterogeneous processing system which allows for intelligent partitioning of a ML model. Unlike prior solutions which rely on user preferences to assign the ML model, the unified interface can violate or break the user preferences when partitioning the ML model. The unified interface can receive information from the processors (e.g., a NPU, CPU, GPU, etc.) and determine the capabilities, current workload, power metrics, subgraphs of the ML model they can execute, and the like. With this information, the unified interface can intelligently choose when to violate or break the user-entered priority based instructions.
A phase-locked loop circuit includes a phase-frequency detector circuit, a time amplifier circuit, a control circuit, and a variable oscillator circuit, such as a voltage-controlled oscillator. The phase-frequency detector is configured to generate a phase error output indicating a phase difference between a reference signal and a feedback signal generated using an output signal. The time amplifier is configured to extend the phase difference to generate an extended phase error output. The variable oscillator is configured to use one or more control signals generated by the control circuit using the extended phase error to adjust a frequency of the output signal. The control circuit may generate a proportional control signal using a switched-resistor circuit that includes a pull-up resistor and a pull-down resistor. The control circuit may generate an integral control signal using a capacitive-shared-integration circuit that includes passive integrating components.
H04L 7/033 - Commande de vitesse ou de phase au moyen des signaux de code reçus, les signaux ne contenant aucune information de synchronisation particulière en utilisant les transitions du signal reçu pour commander la phase de moyens générateurs du signal de synchronisation, p. ex. en utilisant une boucle verrouillée en phase
41.
METHOD FOR SCALABLE OPERATOR LOADING FOR PROGRAM MEMORY LIMITED ACCELERATORS
A method includes allocating, at least one boot image, and parsing programs comprising one or more operators. The method further includes determining a count corresponding to each operator included in each program and an estimated program memory consumption of each operator. The count corresponds to each operator and indicates a quantity of times each operator in each program is detected. The at least one boot image is populated with common operators based on a program memory size limit of an hardware accelerator circuitry of an IC device and an estimated program memory consumption of each common operator. The common operators are operators that are within a threshold count. The at least one boot image is populated with non-common operators based on the program memory size limit and an estimated program memory consumption of each non-common operator. The non-common operators are operators that are not within the threshold count.
Optimizing timing margins across conditions is described. In one or more implementations, a computing system may include an interface circuitry configured to adjust a timing alignment of first and second signals between a central processing unit (CPU) of the system and a device coupled with the CPU and to measure and store one or more margins between the timing alignment and misalignments of the first and second signals. The interface circuitry may be configured to measure and store the timing margins at first and second conditions. The first and second conditions may be different voltages, temperatures, etc. The system may be configured to force the first and/or second condition. The system may be configured to calculate a coefficient from differences in between the margins and between the first and second conditions.
In-system electrical connectivity detection. In one or more implementations, a computing device includes a transmitter and a receiver in a package, the transmitter to transmit a signal to a separate device, the receiver to receive and measure a reflection of the transmitted signal, and the measured reflection for characterizing (e.g., testing or detecting) an electrical connection between the computing and separate devices. The computing device may characterize (e.g., detect a discontinuity in) the electrical connection by comparing a magnitude of the transmitted signal with a magnitude of the measured reflection. The computing device may be coupled with the separate device by multiple electrical connections, and the multiple electrical connections may be tested by corresponding transmitters and receivers.
Fine-grained preemption of a data flow architecture based neural processing unit (NPU) includes executing, by a controller, control-code that implements a first context in the NPU. In response to the controller detecting a preemption opcode in the control-code, detecting, by the controller, a second context awaiting execution by the neural processing unit. The second context has a priority that is greater than a priority of the first context. In response to detecting the second context, the NPU switches from executing the first context to implementing the second context.
G06N 3/063 - Réalisation physique, c.-à-d. mise en œuvre matérielle de réseaux neuronaux, de neurones ou de parties de neurone utilisant des moyens électroniques
45.
HIGH SPEED SERIAL PROTOCOL LOW FREQUENCY PERIODIC SIGNALING DETECTION
A detector for detecting a periodic square wave (PSW) includes: an interface circuit configured to generate a serial bit stream (SBS) by sampling an input signal carrying the PSW; a decimation encoder circuit (DEC) coupled to the interface circuit and configured to generate, for every first number of bits in the SBS, a symbol in a symbol stream; a pulse width estimator (PWE) configured to generate, for each symbol generated by the DEC, an estimate of a width of a most recent pulse (MRP) in a sequence of bits in the SBS, where the sequence of bits correspond to a pre-determined number of most recent symbols generated by the DEC; and a state machine coupled to the PWE and configured to declare detection of the PSW in response to detecting that the estimate of the width of the MRP is within a pre-determined range for a pre-determined period of time.
Embodiments herein describe techniques for setting and using saliency values to modify how a ML model is trained. In one embodiment, different blocks of data (referred to herein as tiles) are assigned respective saliency values. After performing one or more iterations, a training application can modify the default saliency values assigned to the tiles to reflect the importance of the tile. In one embodiment, the training application evaluates a weight gradient that indicates how a weight (or weights) corresponding to each tile are modified and modifies the saliency values accordingly. The ML training system can then use the saliency values to affect future training iterations to reduce the time required to train the ML model or save power.
G06F 18/214 - Génération de motifs d'entraînementProcédés de Bootstrapping, p. ex. ”bagging” ou ”boosting”
G06V 10/46 - Descripteurs pour la forme, descripteurs liés au contour ou aux points, p. ex. transformation de caractéristiques visuelles invariante à l’échelle [SIFT] ou sacs de mots [BoW]Caractéristiques régionales saillantes
G06V 10/75 - Organisation de procédés de l’appariement, p. ex. comparaisons simultanées ou séquentielles des caractéristiques d’images ou de vidéosApproches-approximative-fine, p. ex. approches multi-échellesAppariement de motifs d’image ou de vidéoMesures de proximité dans les espaces de caractéristiques utilisant l’analyse de contexteSélection des dictionnaires
47.
DETERMINISTIC JITTER DETECTION AND MITIGATION FOR FREQUENCY DIVIDER CIRCUITRY
A communication system includes correction circuitry. The correction circuitry includes detection circuitry and clock correction circuitry. The detection circuitry determines a first correction value based on a first rising edge and a second rising edge of a divided clock signal. The clock correction circuitry receives a first clock signal, a second clock signal and the first correction value, and generates a first adjusted clock signal and a second adjusted clock signal based on the first correction value. The first clock signal and the second clock signal. The first adjusted clock signal and the second adjusted clock signal are used by divider circuitry to generate the divided clock signal. A frequency of the divided clock signal is less than a frequency of the first clock signal and the second clock signal.
Fine-grained preemption of a data flow architecture based neural processing unit (NPU) includes executing, by a controller, control-code that implements a first context in the NPU. In response to the controller detecting a preemption opcode in the control-code, detecting, by the controller, a second context awaiting execution by the neural processing unit. The second context has a priority that is greater than a priority of the first context. In response to detecting the second context, the NPU switches from executing the first context to implementing the second context.
Command stream stitching for hardware acceleration includes generating, by a host processor, a stitched block representing a plurality of commands for a hardware accelerator. The host processor generates a stitched command from the plurality of commands. The stitched command references the stitched block. The hardware accelerator executes the stitched block in response to invoking the stitched command. The hardware accelerator generates a single notification directed to the host processor for the stitched command.
Techniques are described for the use of optical couplers with photonic integrated circuits (PICs). Some techniques include optically coupling one device (e.g., a package) to another via one or more optical couplers to, for instance, optically couple an on-chip waveguide to an off-chip optical fiber. Some techniques do not require careful alignment of a waveguide with a lens or other components, and are scalable to large numbers of optical paths. Moreover, some optical couplers herein may be fabricated with known (e.g., conventional) wafer-level processes.
A memory device includes memory cells, driver circuitry, and control circuitry. The driver circuitry is connected to the memory cells. The driver circuitry includes first sense amplifier circuitries connected to first memory cells of the memory cells. The second sense amplifier circuitries are connected to second memory cells of the memory cells. The control circuitry enables the first sense amplifier circuitries and disables the second sense amplifier circuitries based on a read command. First data associated with the first memory cells is output via the first sense amplifier circuitries based on the read command.
A unit cell circuitry for a digital-to-analog converter (DAC) circuitry includes cascode circuitry, switch circuitry, and capacitor circuitry. The cascode circuitry is connected to a first output node and a second output node of the unit cell circuitry. The switch circuitry is connected to the cascode circuitry. The capacitor circuitry includes one or more capacitors connected to the switch circuitry. The switch circuitry connects the capacitor circuitry to the cascode circuitry to inject a charge onto the cascode circuitry. The unit cell circuitry outputs a signal based on the injected charge.
Command stream stitching for hardware acceleration includes generating, by a host processor, a stitched block representing a plurality of commands for a hardware accelerator. The host processor generates a stitched command from the plurality of commands. The stitched command references the stitched block. The hardware accelerator executes the stitched block in response to invoking the stitched command. The hardware accelerator generates a single notification directed to the host processor for the stitched command.
Embodiments herein describe an integrated circuit (IC) including an integrated circuit including a first die and a second die including an inductor and disposed over the first die, where the second die is electrically coupled to the first die via hybrid bonds (MBs). The IC may include a metal layer disposed under a portion of the inductor. The IC may further include a shielding layer disposed within the first die. The IC may also include first metal strips disposed adjacent a head section of the inductor and second metal strips disposed over a leg section of the inductor. The inductor may include a head section constructed as a dual loop and a leg section constructed as a pair of legs, where the inductor is enclosed within isolation walls.
A metastable logic-based physical unclonable function (PUF) circuit with trimmable source degeneration and enhanced aging performance includes metastable logic that generates metastable states at first and second outputs, which settle into respective determinable logic states based on random process variations of elements of the metastable logic. The PUF circuit further includes compensation circuitry to compensate for load mismatches of the first and second outputs. The PUF circuit may further include trimmable (e.g., selectable) source degeneration resistors for controlling voltages of transistors of the metastable logic. Varying numbers of source degeneration resistors may be selected to control the determinable logic state and/or to evaluate the PUF circuit for stability. The source degeneration resistors may also serve a power gates to disable the PUF circuit when not in use, which may reduce age-related effects.
H04L 9/32 - Dispositions pour les communications secrètes ou protégéesProtocoles réseaux de sécurité comprenant des moyens pour vérifier l'identité ou l'autorisation d'un utilisateur du système
G11C 16/04 - Mémoires mortes programmables effaçables programmables électriquement utilisant des transistors à seuil variable, p. ex. FAMOS
56.
INDUCTOR DEGRADATION REDUCTION IN 3D STACKED INTEGRATION WITH HYBRID BOND
Embodiments herein describe an integrated circuit (IC) including an integrated circuit including a first die and a second die including an inductor and disposed over the first die, where the second die is electrically coupled to the first die via hybrid bonds (HBs). The IC may include a metal layer disposed under a portion of the inductor. The IC may further include a shielding layer disposed within the first die. The IC may also include first metal strips disposed adjacent a head section of the inductor and second metal strips disposed over a leg section of the inductor. The inductor may include a head section constructed as a dual loop and a leg section constructed as a pair of legs, where the inductor is enclosed within isolation walls.
H01L 23/00 - Détails de dispositifs à semi-conducteurs ou d'autres dispositifs à l'état solide
H01F 17/00 - Inductances fixes du type pour signaux
H01L 23/48 - Dispositions pour conduire le courant électrique vers le ou hors du corps à l'état solide pendant son fonctionnement, p. ex. fils de connexion ou bornes
H01L 23/522 - Dispositions pour conduire le courant électrique à l'intérieur du dispositif pendant son fonctionnement, d'un composant à un autre comprenant des interconnexions externes formées d'une structure multicouche de couches conductrices et isolantes inséparables du corps semi-conducteur sur lequel elles ont été déposées
H01L 23/552 - Protection contre les radiations, p. ex. la lumière
H01L 23/58 - Dispositions électriques structurelles non prévues ailleurs pour dispositifs semi-conducteurs
H01L 25/16 - Ensembles consistant en une pluralité de dispositifs à semi-conducteurs ou d'autres dispositifs à l'état solide les dispositifs étant de types couverts par plusieurs des sous-classes , , , , ou , p. ex. circuit hybrides
57.
SYSTEMS AND METHODS FOR MANAGING ORDER OF COMMAND PROCESSING
A computer-implemented method for managing processing order for a plurality of commands can include in response to receiving each command of a plurality of commands in a receipt order, assigning each respective command of the plurality of commands to a respective processing queue of a plurality of processing queues to be processed, and setting, for each of the plurality of commands and in the receipt order, an identifier based on the respective queue assigned to each of the plurality of commands, and managing, based on the identifiers for each of the plurality of commands in the receipt order, an order of processing of each of the plurality of commands from the respective processing queue of the plurality of processing queues. Various other methods, systems, and computer-readable media are also disclosed.
Examples herein describe a three-dimensional (3D) die stack. The 3D die stack includes a programmable logic (PL) die and a compute die stacked on top of the PL die. The PL die includes a plurality of configurable blocks and a plurality of first electrical connections on a top side of the PL die. The compute die includes a plurality of data processing engines and a plurality of second electrical connections on a bottom side of the compute die. The three-dimensional die stack includes a plurality of tiles, each tile comprising M configurable blocks included in the plurality of configurable blocks and N data processing engines included in the plurality of data processing engines.
G06F 15/80 - Architectures de calculateurs universels à programmes enregistrés comprenant un ensemble d'unités de traitement à commande commune, p. ex. plusieurs processeurs de données à instruction unique
59.
DYNAMIC OPERATOR DISPATCH MODE FOR IMPLEMENTING MACHINE-LEARNING MODELS
To enable an accelerator unit to perform one or more operators for a machine-learning model, a processing system is configured to generate a launch kernel using a dynamic operator dispatch mode. For example, a processing unit of the processing system first organizes an operator group of the machine-learning model into a series of nodes that represents the operators in the operator group. Based on this series of nodes, the processing unit retrieves and modifies pre-compiled operators from an operator library stored in a memory of the processing system. The processing unit then generates a launch kernel based on the modified pre-compiled operators.
Embodiments herein describe a method including posting a persistent work request (P-WR) to a memory space accessible by a control frontend (CFE), the P-WR including information about data chunks, ringing a doorbell in a control frontend (CFE), allowing the CFE to inspect the P-WR, allowing the CFE to set up virtual doorbells at specific addresses in the memory space, specifying a doorbell address range and completion flags associated with one or a set of data chunks, and allowing the CFE to set up a required data movement for each data chunk and provide completion status.
Accelerated remote-direct-memory-access (RDMA) command construction for GPU-directed fine-grained communication, including interpreter logic that frees a host device and/or compute units of a data processing element (DPE), such as graphic processing unit (GPU), from managing execution of the WRs. The interpreter logic frees the host/DPE from managing execution of the WRs. The interpreter logic may access memory of the host/DPE (e.g., data, work request queues, completion queues, etc.), such as to retrieve the WRs and/or to write completion notifications, and/or the host/DPE may write the WRs to registers accessible to the interpreter logic.
G06F 9/48 - Lancement de programmes Commutation de programmes, p. ex. par interruption
G06F 15/173 - Communication entre processeurs utilisant un réseau d'interconnexion, p. ex. matriciel, de réarrangement, pyramidal, en étoile ou ramifié
62.
QUANTIZING LOW-PRECISION NEURAL NETWORKS FOR LOSSLESS ACCUMULATION
Embodiments herein relate to training, calibrating, or preparing quantized neural networks for lossless low-precision accumulation with floating-point dot product operands. This includes an improvement in training low-precision floating-point neural networks for a certain accumulator bit width. The accumulator-aware weight quantization methodology yields floating-point weights that avoid arithmetic errors caused by overflow, underflow, and rounding, among other common problems when accumulated into low-precision accumulators during dot products with input activations of known data formats. This is implemented by imposing constraints on the values of the exponents and mantissas used to represent the weights of the neurons of the neural network.
G06F 7/544 - Méthodes ou dispositions pour effectuer des calculs en utilisant exclusivement une représentation numérique codée, p. ex. en utilisant une représentation binaire, ternaire, décimale utilisant des dispositifs n'établissant pas de contact, p. ex. tube, dispositif à l'état solideMéthodes ou dispositions pour effectuer des calculs en utilisant exclusivement une représentation numérique codée, p. ex. en utilisant une représentation binaire, ternaire, décimale utilisant des dispositifs non spécifiés pour l'évaluation de fonctions par calcul
G06F 7/556 - Méthodes ou dispositions pour effectuer des calculs en utilisant exclusivement une représentation numérique codée, p. ex. en utilisant une représentation binaire, ternaire, décimale utilisant des dispositifs n'établissant pas de contact, p. ex. tube, dispositif à l'état solideMéthodes ou dispositions pour effectuer des calculs en utilisant exclusivement une représentation numérique codée, p. ex. en utilisant une représentation binaire, ternaire, décimale utilisant des dispositifs non spécifiés pour l'évaluation de fonctions par calcul de fonctions logarithmiques ou exponentielles
Examples herein describe a peripheral I/O device with a hybrid gateway that permits the device to have both I/O and coherent domains. As a result, the compute resources in the coherent domain of the peripheral I/O device can communicate with the host in a similar manner as CPU-to-CPU communication in the host. The dual domains in the peripheral I/O device can be leveraged for machine learning (ML) applications. While an I/O device can be used as an ML accelerator, these accelerators previously only used an I/O domain. In the embodiments herein, compute resources can be split between the I/O domain and the coherent domain where a ML engine is in the I/O domain and a ML model is in the coherent domain. An advantage of doing so is that the ML model can be coherently updated using a reference ML model stored in the host.
Systems and methods for dynamically selecting source code development tests using a trained machine learning (ML) model are disclosed. In certain embodiments, a plurality of data features is derived from information indicating a plurality of modifications to a source code repository. Based at least in part on the derived data features, an ML model is trained to identify correlations between the modifications and a plurality of historical source code development test results. Upon receiving an indication of one or more additional modifications to the source code repository, the trained ML model dynamically selects, from a plurality of source code development tests, a subset of source code development tests relevant to the additional modifications.
A processing unit executes a sequence of instructions comprising a branch instruction and selects a target branch predictor for the branch instruction from a plurality of branch predictors. The selection is based on a language model trained on a set of training instruction sequences that comprise branch instructions. The target branch predictor then determines the outcome of the branch instruction. The language model is trained to identify the branch instructions as having easy-to-predict outcomes or hard-to-predict outcomes based on sequence of instructions. Information generated by the language model is provided to a branch prediction unit, which uses the information to determine whether branch instructions are easy-to-predict or hard-to-predict.
Embodiments herein relate to modifying the framework of an FFN module of a machine learning model. Modifications include an improved nonlinear function of that aims to decrease the number of hidden dimensions of the FFN module, thereby reducing the computational cost.
Computer-implemented technology mapping of a circuit design includes, for a node of a circuit design, generating a plurality of regular cuts and a plurality of super cuts. A regular cut subset that is a subset of the plurality of regular cuts having M highest priorities and a super cut subset that is a subset of the plurality of super cuts having N highest priorities are generated. Each super cut that is incompatible with a cascaded lookup-table (LUT) circuit structure is discarding from the super cut subset. A cut from a group of cuts including the regular cut subset and the super cut subset is selected to implement a portion of the circuit design including the node.
G06F 30/327 - Synthèse logiqueSynthèse de comportement, p. ex. logique de correspondance, langage de description de matériel [HDL] à liste d’interconnections [Netlist], langage de haut niveau à langage de transfert entre registres [RTL] ou liste d’interconnections [Netlist]
G06F 1/03 - Générateurs de fonctions numériques travaillant, au moins partiellement, par consultation de tables
A processing unit executes a sequence of instructions comprising a branch instruction and selects a target branch predictor for the branch instruction from a plurality of branch predictors. The selection is based on a language model trained on a set of training instruction sequences that comprise branch instructions. The target branch predictor then determines the outcome of the branch instruction. The language model is trained to identify the branch instructions as having easy-to-predict outcomes or hard-to-predict outcomes based on sequence of instructions. Information generated by the language model is provided to a branch prediction unit, which uses the information to determine whether branch instructions are easy-to-predict or hard-to-predict.
A circuit arrangement includes a plurality of cryptographic accelerators. Each cryptographic accelerator is configured to perform cryptographic operations according to a respective cryptographic protocol. A first memory is coupled to the cryptographic accelerators. A first processor is configured to specify, in response to requests to perform the cryptographic operations, parameters to the cryptographic accelerators according to the requests. The first processor is configured to identify, in the first memory, keys that are associated with the cryptographic accelerators, and signal the cryptographic accelerators to commence performing the cryptographic operations according to the parameters and using the associated keys.
H04L 9/06 - Dispositions pour les communications secrètes ou protégéesProtocoles réseaux de sécurité l'appareil de chiffrement utilisant des registres à décalage ou des mémoires pour le codage par blocs, p. ex. système DES
H04L 9/30 - Clé publique, c.-à-d. l'algorithme de chiffrement étant impossible à inverser par ordinateur et les clés de chiffrement des utilisateurs n'exigeant pas le secret
G06F 13/28 - Gestion de demandes d'interconnexion ou de transfert pour l'accès au bus d'entrée/sortie utilisant le transfert par rafale, p. ex. acces direct à la mémoire, vol de cycle
70.
OFFLOADING OPERATIONS USING A NETWORK INTERFACE CONTROLLER
Offloading operations for a computing system includes executing an application by a Central Processing Unit (CPU) of the computing system. The application includes a first set of operations and a second set of operations. The first set of operations may be executed by a Graphics Processing Unit of the computing system. The Graphics Processing Unit may execute the first set of operations under the control of the CPU. The second set of operations may be executed by a Smart Network Interface Controller of the computing system. The Smart Network Interface Controller may execute the second set of operations under control of the CPU.
Some examples described herein provide for interconnect in chiplet systems, for example system-level techniques for error correction in chip-to-chip interfaces. In an example, a method of error correction includes receiving, at a first chiplet, a data message via a set of interconnect, and transmitting a first control message that requests retransmission of the data message based on detecting an error associated with receiving the data message. The method also includes transmitting one or more instances of a second control message that indicates an idle operation at the first chiplet until the first chiplet receives a third control message that triggers an end of a retransmission mode. The method also includes transmitting a fourth control message frame indicating the end of the retransmission mode, and receiving a retransmission of the data message from the second chiplet.
Automatic run-time skew-aware optimization of rooted collectives includes determining skews amongst compute nodes of a distributed computing system, as an application executes on the compute nodes, and determining implementations for collective operations of the application based at least in part on the skews. Skews may be determined based on timestamps of operations of the application program. Time stamps of one or more of the compute nodes may be estimated or inferred from timestamps of other compute nodes. Timestamps may be aggregated to determine global skews. Collective implementations may be determined for a sequence of collective operations based on skew impacts amongst the sequence of collective operations. Subsequent collective operations may be predicted based on current collective operations and a history of persistent collective operations, and implementations may be determined for the predicted collective operations prior to receipt of calls for the predicted collective operations.
An emulation system includes a first platform device, a second platform device, and a processing device. The first platform device includes first integrated circuit (IC) devices. The first IC devices emulate a circuit design for verification of the circuit design. The second platform device includes second IC devices. The processing device connected to the first platform device and the second platform device. The processing device configured to emulate the circuit design using one or more of the second IC devices based on a failure associated with the first platform device being detected.
G06F 30/33 - Vérification de la conception, p. ex. simulation fonctionnelle ou vérification du modèle
G06F 9/455 - ÉmulationInterprétationSimulation de logiciel, p. ex. virtualisation ou émulation des moteurs d’exécution d’applications ou de systèmes d’exploitation
74.
OFFLOADING OPERATIONS USING A NETWORK INTERFACE CONTROLLER
Offloading operations for a computing system includes executing an application by a Central Processing Unit (CPU) of the computing system. The application includes a first set of operations and a second set of operations. The first set of operations may be executed by a Graphics Processing Unit of the computing system. The Graphics Processing Unit may execute the first set of operations under the control of the CPU. The second set of operations may be executed by a Smart Network Interface Controller of the computing system. The Smart Network Interface Controller may execute the second set of operations under control of the CPU.
Data transfer over a packet-based network-on-chip (NoC) of an integrated circuit device, including an example in which a first region of programmable logic (PL) serves as a first interface circuit between a first circuit block and a NoC master unit (NMU), to receive first and second data via respective first and second channels of the first circuit block based on a communication protocol of the first circuit block, concatenate the first and second data to provide the concatenated content, and transmit the concatenated content to the NMU. The NoC may route the packets from the NMU to a NoC slave unit (NSU) associated with a second circuit block via a pre-determined route of the NoC that is dedicated to traffic between the first and second circuit blocks. A second region of the PL serves as an interface circuit between the NSU and the second block to unpack the data.
H04L 49/109 - Éléments de commutation de paquets caractérisés par la construction de la matrice de commutation intégrés sur micropuce, p. ex. interrupteurs sur puce
H04L 45/00 - Routage ou recherche de routes de paquets dans les réseaux de commutation de données
Configuring a System-on-Chip (SoC) can include receiving first configuration data for an SoC. The SoC has a plurality of hardware systems including programmable logic. The first configuration data includes a portion specifying a circuit design for the programmable logic. A user input specifying a change to a non-netlist configuration setting of the SoC can be received. Second configuration data for the SoC can be generated. The second configuration data specifies the change to the non-netlist configuration setting. Updated configuration data for the SoC can be generated by merging the second configuration data with the first configuration data leaving the portion of the first configuration data specifying the circuit design unchanged.
A circuit arrangement includes a plurality of cryptographic accelerators. Each cryptographic accelerator is configured to perform cryptographic operations according to a respective cryptographic protocol. A first memory is coupled to the cryptographic accelerators. A first processor is configured to specify, in response to requests to perform the cryptographic operations, parameters to the cryptographic accelerators according to the requests. The first processor is configured to identify, in the first memory, keys that are associated with the cryptographic accelerators, and signal the cryptographic accelerators to commence performing the cryptographic operations according to the parameters and using the associated keys.
H04L 9/32 - Dispositions pour les communications secrètes ou protégéesProtocoles réseaux de sécurité comprenant des moyens pour vérifier l'identité ou l'autorisation d'un utilisateur du système
78.
CARRY CHAIN REDUCTION AND REPLACEMENT FOR CIRCUIT DESIGNS
Carry chain reduction and/or replacement for circuit designs includes detecting, by computer hardware, a carry chain instance of a circuit design that is under-utilized. The carry chain instance has a first number of logic levels and includes one or more carry chain primitives dedicated for implementing carry chain logic. An updated carry chain instance is generated by the computer hardware. The updated carry chain instance has a second number of logic levels less than the first number of logic levels. The carry chain instance is replaced with the updated carry chain instance within the circuit design by the computer hardware.
Embodiments herein describe a system including a plurality of hardware accelerators including at least one mixture-of-experts (MoE) layer having multiple experts and a plurality of network interface cards (NICs) coupled to the plurality of hardware accelerators, wherein at least one expert of the multiple experts is offloaded from the plurality of hardware accelerators to the plurality of NICs. The plurality of hardware accelerators may be graphics processing units (GPUs). In one example, a subset of the multiple experts are selectively offloaded from the plurality of GPUs to the plurality of NICs based on memory and computational capacity available on the plurality of NICs. In another example, the multiple experts are designated as either hot experts or cold experts. The cold experts are offloaded from the plurality of GPUs to the plurality of NICs and the hot experts are duplicated for each of the plurality of GPUs.
Chip packages are described herein that includes integrated passive devices embedded in a core of a substrate of the chip package, such as a package substrate or an interposer, that shield routings coupled to inductors from adjacent through-substrate conductive paths (e.g., vias). In one example, a chip package includes an integrated circuit (IC) die mounted to a substrate. A core of the substrate has a plurality of inductor routing vias, a plurality of signal transmission vias, and a plurality of ground and power routing vias. A first integrated passive device (IPD) is disposed in the core and separates at least one of the plurality of inductor routing vias from an adjacent via, the adjacent via being one of the plurality of signal transmission vias or one of the plurality of ground and power routing vias.
H01L 23/31 - Encapsulations, p. ex. couches d’encapsulation, revêtements caractérisées par leur disposition
H01L 25/00 - Ensembles consistant en une pluralité de dispositifs à semi-conducteurs ou d'autres dispositifs à l'état solide
H01L 25/16 - Ensembles consistant en une pluralité de dispositifs à semi-conducteurs ou d'autres dispositifs à l'état solide les dispositifs étant de types couverts par plusieurs des sous-classes , , , , ou , p. ex. circuit hybrides
81.
ROBUST FLAT-TOP OPTICAL FILTER DESIGN UTILIZING PHASE AND POWER SPLITTING RATIO OF MMI COUPLERS
A filter circuitry includes a first multi-mode interferometer (MMI) circuitry receives an optical input signal and generates a first output optical signal and a second output optical signal according to a first power splitting ratio. A second MMI circuitry receives the first output optical signal and the second output optical signal and generates a third output optical signal and a fourth output optical signal according to a second power splitting ratio. A third MMI circuitry receives the third output optical signal and the fourth output optical signal and generates a fifth output optical signal and a sixth output optical signal according to the second power splitting ratio. A fourth MMI circuitry receives the fourth output optical signal and the fifth output optical signal and generates a seventh output optical signal and an eighth output optical signal according to a third power splitting ratio, with each power splitting ratio being different.
H04B 10/079 - Dispositions pour la surveillance ou le test de systèmes de transmissionDispositions pour la mesure des défauts de systèmes de transmission utilisant un signal en service utilisant des mesures du signal de données
H04J 14/02 - Systèmes multiplex à division de longueur d'onde
82.
ROBUST SINGLE EVENT UPSET (SEU) TOLERANT HIGH-PERFORMANCE FLIP-FLOP
Embodiments herein describe single event upset (SEU) tolerant flip-flop that includes master latch circuitry, slave latch circuitry, and a tristate driver having an input coupled to an output of the master latch circuitry and an output coupled to a first data input of the slave latch circuitry, where the first tristate driver is configured to inhibit charge transfer from the first data input of the slave latch circuitry to the output of the master latch circuitry.
H03K 3/3562 - Circuits bistables du type primaire-secondaire
H03K 19/08 - Circuits logiques, c.-à-d. ayant au moins deux entrées agissant sur une sortieCircuits d'inversion utilisant des éléments spécifiés utilisant des dispositifs à semi-conducteurs
H03K 19/094 - Circuits logiques, c.-à-d. ayant au moins deux entrées agissant sur une sortieCircuits d'inversion utilisant des éléments spécifiés utilisant des dispositifs à semi-conducteurs utilisant des transistors à effet de champ
Clock skew measurement circuitry determines skew between clock signals within an integrated circuit (IC) device or between multiple IC devices. An IC device includes clock tree circuitry and the clock skew measurement circuitry. The clock tree circuitry provides a first clock signal and a second clock signal to components of the IC device. The clock skew measurement circuitry is connected to the clock tree circuitry. The clock skew measurement circuitry generates and outputs an error signal based on a phase difference between the first clock signal and the second clock signal.
H03K 5/135 - Dispositions ayant une sortie unique et transformant les signaux d'entrée en impulsions délivrées à des intervalles de temps désirés par l'utilisation de signaux de référence de temps, p. ex. des signaux d'horloge
Embodiments herein describe generating non-linear thresholds which can be used, for example, in a quantization operation of a neural network layer. In one example, a compiler can receive a set of non-linear thresholds (e.g., rounded integer thresholds) and determine delta values indicating respective differences between two thresholds of the non-linear thresholds. Because the thresholds are non-linear, the delta values are also different. The compiler can calculate residual errors to represent the differences between the deltas with respect to a step size. Using these residual errors, an initial threshold value, a step size between the thresholds, the non-linear thresholds can be calculated as needed, instead of having to store the non-linear thresholds in memory.
G06F 7/60 - Méthodes ou dispositions pour effectuer des calculs en utilisant une représentation numérique non codée, c.-à-d. une représentation de nombres sans baseDispositifs de calcul utilisant une combinaison de représentations de nombres codées et non codées
G06F 5/01 - Procédés ou dispositions pour la conversion de données, sans modification de l'ordre ou du contenu des données maniées pour le décalage, p. ex. la justification, le changement d'échelle, la normalisation
Embodiments herein describe a content adaptive array that can include different types of data. A compute unit can include conversion circuitry (e.g., upcast circuitry) that can identify the datatype(s) in the content adaptive array and convert the data so it has a desired datatype. For example, if the content adaptive array has both FP and INT, the upcast circuitry converts the data into the same datatype (e.g., FP8). If the array includes FP4 and FP8 (or INT4 and INT8), the upcast circuitry converts the data into FP8. This means the circuitry in the compute unit that performs the data operation (e.g., matrix multiplication) does not have to support many different types of datatypes.
A compression-defined broaching mount for compression-attached memory module (CAMM) includes a plurality of pins configured to retain a plurality of circuit boards in a compressed position. The pins may include a snap-pin, a push-pin, and/or a spring-loaded pin. The pin is removable. The CAMM may include a flat washer and/or a counter-sunk washer configured to receive the pin. The counter-sunk washer may have an opening to a chamber therein, where an inner surface of the chamber has a channel to receive a locking member of the pin. The CAMM may include a compressible material to exert a force when the pin is in a locked position. The compressible material may include an O-ring, a gasket, and/or a film. The CAMM may include a visual compression indicator.
A 3D device includes a first semiconductor chip and a second semiconductor chip stacked vertically. The first semiconductor chip includes a first plurality of tiles. The second semiconductor chip includes a second plurality of tiles. A bus electrically couples each of the first plurality of tiles to a corresponding one of the second plurality of tiles based on assignments of the first plurality of tiles and the second plurality of tiles to tile-to-tile pairs that define a minimized sum of bus delays among each possible tile-to-tile pairs. In each tile-to-tile pair, a net electrically couples each of a first plurality of pins to a corresponding one of a second plurality of pins based on assignments of the first plurality of pins to the second plurality of pins that define a minimized sum of net delays among each possible pin-to-pin pairs.
A compression-defined broaching mount for compression-attached memory module (CAMM) includes a plurality of pins configured to retain a plurality of circuit boards in a compressed position. The pins may include a snap-pin, a push-pin, and/or a spring-loaded pin. The pin is removable. The CAMM may include a flat washer and/or a counter-sunk washer configured to receive the pin. The counter-sunk washer may have an opening to a chamber therein, where an inner surface of the chamber has a channel to receive a locking member of the pin. The CAMM may include a compressible material to exert a force when the pin is in a locked position. The compressible material may include an O-ring, a gasket, and/or a film. The CAMM may include a visual compression indicator.
Embodiments herein describe techniques for spatially unrolling thresholds (e.g., steps) for traversing a binary search tree. A binary search tree permits quantization logic to quickly search through thresholds stored in memory to perform quantization (e.g., convert an input value into one of the thresholds). Assuming the thresholds are sorted in order when stored in memory, the result of comparing the input value to a threshold in the current level of the binary tree can be used to select the address of the next threshold in the next level of the binary tree. This permits the quantization logic to traverse the thresholds in a logarithmic manner.
Systems and techniques for selectable slice mapping in shared cache levels are described. In one example, a processor includes a cache system having a shared cache level of a hierarchy of cache levels and slice hashing circuitry associated with the shared cache level. The shared cache level includes multiple slices accessible by threads running on multiple processor cores. The slice hashing circuitry assigns memory addresses used by a particular thread to a subset of the multiple slices closest to the processor core on which the thread runs. The assignment of the slice subset is based on the latency requirements or the data usage of the thread in at least one implementation. The described techniques improve tail latencies for multiple core systems and alleviate the need for additional interconnections for shared cache levels.
A scalar processor associated with a vector processor reduces the quantization error for blocked data with a relatively small register size by predicting adjustments for shared scalars used in runtime quantization. The scalar processor provides a recommended scale value to the vector processor for scaling a block of data from a wide data type format to a narrow data type format. The scalar processor and the vector processor share a register at which the scalar processor stores the recommended scale value and from which the vector processor accesses the recommended scale value. The vector processor performs an operation to quantize at least a portion of the block of data by applying a scale value that is based on the recommended scale value.
G06F 15/80 - Architectures de calculateurs universels à programmes enregistrés comprenant un ensemble d'unités de traitement à commande commune, p. ex. plusieurs processeurs de données à instruction unique
Embodiments herein describe a calibration circuit including a multi-phase clock generator configured to receive a clock signal and generate a multi-phase clock output, the multi-phase clock generator including a first injection-locked oscillator, a mixer, a current digital-to-analog converter (IDAC), and a voltage to current converter (VTOI), the IDAC and VTOI configured to fine tune offsets of the mixer and an input of the VTOI, and overcome phase errors from injection locking disturbance. The multi-phase clock generator further includes a pre-skew buffer configured to receive the multi-phase clock output from the multi-phase clock generator and generate multiple data signals and a phase interpolator (PI) configured to receive the multiple data signals from the pre-skew buffer and generate a shifted clock signal of the clock signal received by the multi-phase clock generator.
A scalar processor (142) associated with a vector processor (140) reduces the quantization error for blocked data with a relatively small register size by predicting adjustments for shared scalars used in runtime quantization. The scalar processor provides a recommended scale value (210) to the vector processor for scaling a block of data from a wide data type format to a narrow data type format. The scalar processor and the vector processor share a register (206) at which the scalar processor stores the recommended scale value and from which the vector processor accesses the recommended scale value. The vector processor performs an operation to quantize at least a portion of the block of data (412) by applying a scale value that is based on the recommended scale value.
G06F 15/80 - Architectures de calculateurs universels à programmes enregistrés comprenant un ensemble d'unités de traitement à commande commune, p. ex. plusieurs processeurs de données à instruction unique
G06F 9/30 - Dispositions pour exécuter des instructions machines, p. ex. décodage d'instructions
G06F 9/38 - Exécution simultanée d'instructions, p. ex. pipeline ou lecture en mémoire
Implementing a circuit design for an integrated circuit device having a plurality of dies where each die has a plurality of fabric sub-regions (FSRs) includes detecting an inter-FSR net of the circuit design. A source of the inter-FSR net and an anchor for the inter-FSR net are projected to a selected die of the plurality of dies resulting in a projected source and a projected anchor in the selected die. The inter-FSR net is replaced with a modified inter-FSR net coupling the projected source with the projected anchor in the selected die and a new intra-FSR net coupling the anchor with a load of the inter-FSR net. The circuit design including the modified inter-FSR net and the new intra-FSR net is routed.
Embodiments herein relate to implementing a downsampling technique that uses a non integer stride, within the architecture of a vision transformer. Furthermore, embodiments herein relate to implementing a masked auto-encoder architecture to facilitate training the flexible, non integer stride downsampling layer. This reduces computational costs while increasing classification performance.
A photonics device, a co-packaged device, and a computer system that provide wide spectral range optical coupling and methods for using the same are disclosed herein. In one example, a photonics device is provided that includes a device and dielectric layer, and a first coupling device. The device and dielectric layer includes a first recess within a first surface of the device and dielectric layer, and a second recess within the first surface of the device and dielectric layer. The first coupling device is mounted to the first surface of the device and dielectric layer. The first coupling device includes a first mounting element and a first extended region. The first mounting element is disposed within the first recess and the first extended region is disposed within the second recess. The first extended region is configured to reflect an optical signal through the device and dielectric layer.
G02B 6/42 - Couplage de guides de lumière avec des éléments opto-électroniques
H01L 23/00 - Détails de dispositifs à semi-conducteurs ou d'autres dispositifs à l'état solide
H01L 25/16 - Ensembles consistant en une pluralité de dispositifs à semi-conducteurs ou d'autres dispositifs à l'état solide les dispositifs étant de types couverts par plusieurs des sous-classes , , , , ou , p. ex. circuit hybrides
97.
EDGE COUPLING TO SURFACE COUPLING PHOTONICS DEVICE
A photonics device, a co-packaged device, and a computer system that provide wide spectral range optical coupling and methods for using the same are disclosed herein. In one example, a photonics device is provided that includes a device and dielectric layer, and a first coupling device. The device and dielectric layer includes a first recess within a first surface of the device and dielectric layer, and a second recess within the first surface of the device and dielectric layer. The first coupling device is mounted to the first surface of the device and dielectric layer. The first coupling device includes a first mounting element and a first extended region. The first mounting element is disposed within the first recess and the first extended region is disposed within the second recess. The first extended region is configured to reflect an optical signal through the device and dielectric layer.
G02B 6/42 - Couplage de guides de lumière avec des éléments opto-électroniques
G02B 6/12 - Guides de lumièreDétails de structure de dispositions comprenant des guides de lumière et d'autres éléments optiques, p. ex. des moyens de couplage du type guide d'ondes optiques du genre à circuit intégré
98.
JTAG-based apparatus and method for input clock frequency measurement
In an implementation, a method may include receiving an input clock at an input pin of a device, the device having a Joint Test Action Group (JTAG) test access port (TAP). The method may also include counting a first number of cycles of a reference clock using a first counter. The method may furthermore include simultaneously counting a second number of cycles of the input clock received at the input pin using a second counter. The method may in addition include calculating a frequency of the input clock based on the counted first number of cycles of the reference clock, the counted second number of cycles of the input clock, and a frequency of the reference clock.
A method for card retention feature retrofit can include inserting an end of a first portion of a latch into a slot cut in an edge of a first printed circuit board, wherein the first portion of the latch is arranged to engage with a second portion of the latch on a second printed circuit board, and coupling the first portion of the latch to the first printed circuit board at least in part by aligning one or more first attachment points of a clamp coupled to the end of the first portion of the latch with one or more second attachment points of the first printed circuit board. Various other methods and systems are also disclosed.
Examples herein describe revocable cryptographic keys. An integrated circuit includes an input/output interface configured to receive inputs including plaintext user keys, metadata, and revocation bits. Cryptographic circuitry is configured to read a key from a first memory. Plaintext user keys are encrypted based on the key to provide encrypted user keys. Metadata is encrypted based on the key to provide encrypted metadata. Revocation bits are encrypted based on the key to provide encrypted revocation bits. A Galois/Counter Mode (GCM) tag is computed based on the key. A processor is configured to write the encrypted user keys, the encrypted metadata, the encrypted revocation bits, and the GCM tag to a second memory to provision the plaintext user keys.