# NeuroBench: A Framework for Benchmarking Neuromorphic Computing Algorithms and Systems
**Authors**: Jason Yik, Korneel Van den Berghe, Douwe den Blanken, Younes Bouhadjar, Maxime Fabre, Paul Hueber, Weijie Ke, Mina A Khoei, Denis Kleyko, Noah Pacik-Nelson, Alessandro Pierro, Philipp Stratmann, Pao-Sheng Vincent Sun, Guangzhi Tang, Shenqi Wang, Biyan Zhou, Soikat Hasan Ahmed, George Vathakkattil Joseph, Benedetto Leto, Aurora Micheli, Anurag Kumar Mishra, Gregor Lenz, Tao Sun, Zergham Ahmed, Mahmoud Akl, Brian Anderson, Andreas G. Andreou, Chiara Bartolozzi, Arindam Basu, Petrut Bogdan, Sander Bohte, Sonia Buckley, Gert Cauwenberghs, Elisabetta Chicca, Federico Corradi, Guido de Croon, Andreea Danielescu, Anurag Daram, Mike Davies, Yigit Demirag, Jason Eshraghian, Tobias Fischer, Jeremy Forest, Vittorio Fra, Steve Furber, P. Michael Furlong, William Gilpin, Aditya Gilra, Hector A. Gonzalez, Giacomo Indiveri, Siddharth Joshi, Vedant Karia, Lyes Khacef, James C. Knight, Laura Kriener, Rajkumar Kubendran, Dhireesha Kudithipudi, Shih-Chii Liu, Yao-Hong Liu, Haoyuan Ma, Rajit Manohar, Josep Maria Margarit-Taulé, Christian Mayr, Konstantinos Michmizos, Dylan R. Muir, Emre Neftci, Thomas Nowotny, Fabrizio Ottati, Ayca Ozcelikkale, Priyadarshini Panda, Jongkil Park, Melika Payvand, Christian Pehle, Mihai A. Petrovici, Christoph Posch, Alpha Renner, Yulia Sandamirskaya, Clemens JS Schaefer, André van Schaik, Johannes Schemmel, Samuel Schmidgall, Catherine Schuman, Jae-sun Seo, Sadique Sheik, Sumit Bam Shrestha, Manolis Sifalakis, Amos Sironi, Kenneth Stewart, Matthew Stewart, Terrence C. Stewart, Jonathan Timcheck, Nergis Tömen, Gianvito Urgese, Marian Verhelst, Craig M. Vineyard, Bernhard Vogginger, Amirreza Yousefzadeh, Fatima Tuz Zohora, Charlotte Frenkel, Vijay Janapa Reddi
> Harvard University Correspondence to:
> Harvard University Delft University of Technology
> Delft University of Technology
> Forschungszentrum Jülich
> University of Groningen
> Delft University of Technology imec
> SynSense
> Örebro University RISE Research Institutes of Sweden
> Accenture Labs
> Intel Corporation, Intel Labs
> City University of Hong Kong
> imec Eindhoven University of Technology
> Innatera Nanosystems B.V.
> Politecnico di Torino
> NeuroBus
> Centrum Wiskunde & Informatica
> Harvard University
> SpiNNcloud Systems GmbH
> Johns Hopkins University
> Istituto Italiano di Tecnologia
> National Institute of Standards and Technology
> Eindhoven University of Technology
> University of Zurich ETH Zurich Google, Paradigms of Intelligence Team
> Queensland University of Technology
> Cornell University
> University of Manchester
> U Waterloo
> University of Texas at Austin Medici Therapeutics
> University of Zurich ETH Zurich
> University of Notre Dame
> Sony Semiconductor Solutions Europe Sony Europe B.V.
> University of Sussex
> University of Bern University of Zurich ETH Zurich
> University of Pittsburgh
> CentraleSupélec, Université Paris-Saclay
> Yale University
> Instituto de Microelectrónica de Barcelona, IMB-CNM (CSIC)
> Technische Universität Dresden ScaDS.AI Dresden/Leipzig
> Rutgers University
> Forschungszentrum Jülich RWTH Aachen
> Uppsala University
> Korea Institute of Science and Technology
> Heidelberg University
> University of Bern
> Prophesee
> Intel Corporation, Intel Labs ZHAW
> Western Sydney University
> University of Tennessee
> Cornell Tech
> UCI Forschungszentrum Jülich
> National Research Council Canada
> KU Leuven imec
> Sandia National Laboratories
> Technische Universität Dresden
> Delft University of Technology These authors jointly supervised this work
> Harvard University These authors jointly supervised this work
## Abstract
Neuromorphic computing shows promise for advancing computing efficiency and capabilities of AI applications using brain-inspired principles. However, the neuromorphic research field currently lacks standardized benchmarks, making it difficult to accurately measure technological advancements, compare performance with conventional methods, and identify promising future research directions. Prior neuromorphic computing benchmark efforts have not seen widespread adoption due to a lack of inclusive, actionable, and iterative benchmark design and guidelines. To address these shortcomings, we present NeuroBench: a benchmark framework for neuromorphic computing algorithms and systems. NeuroBench is a collaboratively-designed effort from an open community of researchers across industry and academia, aiming to provide a representative structure for standardizing the evaluation of neuromorphic approaches. The NeuroBench framework introduces a common set of tools and systematic methodology for inclusive benchmark measurement, delivering an objective reference framework for quantifying neuromorphic approaches in both hardware-independent (algorithm track) and hardware-dependent (system track) settings. In this article, we outline tasks and guidelines for benchmarks across multiple application domains, and present initial performance baselines across neuromorphic and conventional approaches for both benchmark tracks. NeuroBench is intended to continually expand its benchmarks and features to foster and track the progress made by the research community.
keywords: benchmark, neuromorphic
## Introduction
In recent years, the rapid growth of artificial intelligence (AI) and machine learning (ML) has resulted in increasingly complex and large models in pursuit of higher accuracy and range of use cases [1]. The substantial growth rate of model computation exceeds efficiency gains realized through Moore and Dennard technology scaling [2], indicating a looming limit to continued advancements with existing techniques. This issue is compounded by the open challenges of adapting such methods for resource-constrained edge devices (tinyML) in order to enable pervasive and decentralized intelligence through the Internet of Things (IoT) [3]. As such, the urgency for exploring new resource-efficient and scalable computing architectures has intensified.
Neuromorphic computing has emerged as a promising area in addressing these challenges, aiming to unlock key hallmarks of biological intelligence by porting primitives and computational strategies employed in the brain into engineered computing devices and algorithms [4, 5, 6]. Neuromorphic systems hold a critical position in the investigation of novel architectures, as the brain exemplifies an exceptional model for accomplishing scalable, energy-efficient, and real-time embodied computation.
Initially, the term “neuromorphic” referred specifically to approaches that aimed to emulate the biophysics of the brain by leveraging physical properties of silicon, as proposed by Mead in the 1980’s [7]. However, the field of neuromorphic computing research has since grown to encompass a wide range of brain-inspired computing techniques at the algorithmic, hardware, and system levels [4]. While the range of approaches is diverse, neuromorphic computing research generally utilizes mechanisms emulating or simulating biophysical properties more closely than conventional methods, aiming to reproduce high-level performance and efficiency characteristics of biological neural systems.
Neuromorphic algorithms [8] encompass neuroscience-inspired methods which strive towards goals of expanded learning capabilities, such as predictive intelligence, data efficiency, and adaptation, and include approaches such as spiking neural networks (SNNs) and primitives of neuron dynamics, plastic synapses, and heterogeneous network architectures. Algorithm exploration often makes use of simulated execution on readily-available conventional hardware such as CPUs and GPUs, with the goal of driving design requirements for next-generation neuromorphic hardware.
Neuromorphic systems [9] are composed of algorithms deployed to hardware, which seek greater energy efficiency, real-time processing capabilities, and resilience compared to conventional systems. Neuromorphic hardware utilizes a variety of biologically-inspired hardware approaches, including analog neuron emulation, event-based computation, non-von-Neumann architectures, and in-memory processing. Neuromorphic systems target a wide range of applications, from neuroscientific exploration, to low-power edge intelligence and datacenter-scale acceleration.
Despite its promises, progress in the field of neuromorphic research is impeded due to the absence of fair and widely-adopted objective metrics and benchmarks [10, 8]. Without such benchmarks, the validity of neuromorphic solutions cannot be directly quantified, hindering the research community from measuring technological advancement. Standard and rigorous benchmarking is necessary for the neuromorphic community to objectively assess and compare the achievements of novel approaches, and make evidence-based decisions on which directions show promise for achieving breakthrough efficiency, speed, and intelligence, thereby helping to focus research and commercialization efforts on techniques that concretely improve on prior work and conventional computing. Neuromorphic benchmarks have been previously proposed for classical vision [11, 12] and audition tasks [13], open-loop [14] and closed-loop [15] tasks, and for SNN simulator performance assessment [16]. While prior works have made valuable contributions, there are opportunities to further advance the field by addressing three outstanding challenges:
- Lack of a formal definition. The variety of approaches to exploring brain-inspired principles creates difficulties in defining a set of criteria for what should be benchmarked as a “neuromorphic” solution. Closed definitions can impose narrow assumptions and thus risk unfairly excluding promising methods. This challenge necessitates inclusive benchmarks that can be applied generally across the spectrum of potential approaches, allowing for flexible implementation while focusing on task capabilities and metrics of interest such as temporal processing and efficiency. Furthermore, the benchmarks should ideally allow for direct comparison of neuromorphic and conventional approaches.
- Implementation diversity. A wide array of different frameworks targeting different goals, such as neuroscientific exploration [17] and automatic SNN training [18], are used in neuromorphic research. This diversity, which has been instrumental in exploring the landscape of bio-inspired techniques following different methodologies and abstraction levels, comes at the cost of portability and standardization, which in turn limits the ease of benchmark implementation. Benchmarks require common infrastructure that unites tooling to enable actionable implementation and comparison of new methods.
- Rapid research evolution. Neuromorphic approaches are continually and rapidly evolving as part of an emerging field. As the research community continues to make technological progress, so too should benchmark suites and methodology expand to foster inclusion and capture salient performance metrics. An iterative benchmark framework with structured versioning will facilitate productive foundational and evolving performance evaluation.
<details>
<summary>x1.png Details</summary>

### Visual Description
## Diagram: Algorithm Track vs. System Track
### Overview
The image is a comparative flow diagram illustrating two distinct methodologies for development: the "Algorithm Track" and the "System Track." It contrasts a traditional, isolated algorithmic development process with a holistic "co-innovation" approach that integrates hardware considerations. The diagram highlights how system-level constraints and hardware awareness influence the final metrics.
### Components/Axes
The diagram is organized into two horizontal rows (tracks) and a central processing region.
* **Left Column (Inputs):**
* **Algorithm Track:** Labeled "Algorithm Track" (left of the dataset box). Contains a gray box labeled "Dataset."
* **System Track:** Labeled "System Track" (left of the dataset box). Contains a gray box labeled "Dataset."
* **Center Region (Processing - enclosed in a light gray container):**
* **Top (Algorithm Track):** An orange box labeled "Algorithm."
* **Bottom (System Track):** A larger orange box labeled "Algorithm + Hardware."
* **Right Column (Outputs):**
* **Top (Algorithm Track):** A gray box labeled "Algorithm Metrics."
* **Bottom (System Track):** A gray box labeled "System Metrics."
* **Connectors:**
* **Solid Arrows:** Indicate the primary forward flow of data from Dataset → Processing → Metrics.
* **Dashed Arrows:** Indicate cross-track influence and interaction.
### Detailed Analysis
**1. The Algorithm Track (Top Row)**
* **Flow:** The process begins with a "Dataset," which flows horizontally to the right into the "Algorithm" processing block. The output of this block flows horizontally to the right into "Algorithm Metrics."
* **Nature:** This represents a linear, software-centric development path where the algorithm is developed independently of the underlying hardware.
**2. The System Track (Bottom Row)**
* **Flow:** The process begins with a "Dataset," which flows horizontally to the right into the "Algorithm + Hardware" processing block. The output of this block flows horizontally to the right into "System Metrics."
* **Nature:** This represents a co-design approach where the algorithm and hardware are developed in tandem.
**3. Cross-Track Interactions (Vertical/Dashed)**
* **Algorithm-System Co-Innovation:** A vertical, bidirectional dashed arrow connects the "Algorithm" block (top) and the "Algorithm + Hardware" block (bottom). This indicates a feedback loop or collaborative design process between the two tracks.
* **System-Informed Complexity Metrics:** A vertical, upward-pointing dashed arrow connects the "System Metrics" block (bottom) to the "Algorithm Metrics" block (top). This indicates that the metrics derived from the system track are used to inform or refine the metrics of the algorithm track.
### Key Observations
* **Isolation vs. Integration:** The "Algorithm Track" is visually isolated from hardware considerations, whereas the "System Track" explicitly merges them.
* **Directional Influence:** The "System-Informed Complexity Metrics" arrow points *upward*, suggesting that system-level data is a prerequisite or a corrective factor for standard algorithmic metrics.
* **Co-Innovation:** The bidirectional arrow between the central processing blocks suggests that the "Algorithm" track is not entirely independent; it is influenced by the "Algorithm + Hardware" track, implying that modern algorithm design is increasingly dependent on hardware capabilities.
### Interpretation
This diagram serves as a conceptual framework for **Hardware-Software Co-Design**.
* **The Problem:** Traditional algorithmic research often optimizes for abstract metrics (like accuracy or theoretical FLOPs) without considering the physical constraints of the hardware (memory bandwidth, power consumption, latency). This is represented by the top track.
* **The Solution:** The "System Track" proposes that by designing the algorithm and hardware simultaneously, one can achieve better performance.
* **The "System-Informed" Insight:** The most critical element is the "System-Informed Complexity Metrics" arrow. It suggests that "Algorithm Metrics" are often incomplete or misleading when viewed in a vacuum. By feeding system-level metrics back into the algorithm track, researchers can create more realistic, efficient, and deployable solutions. This diagram argues that the "Algorithm Track" is insufficient on its own for modern, high-performance computing or AI applications.
</details>
Figure 1: The two NeuroBench tracks: algorithms and systems. Grey boxes designate what is defined by the benchmark, and orange boxes indicate what is unique to each solution. Connecting arrows between the two tracks denote the co-innovation between the tracks and the cross-stack innovation enabled by this approach. Between algorithm and system solutions, best-performing results from each track can motivate future solutions to the other. In addition, system metrics and results can inform hardware-independent algorithmic complexity metrics.
To tackle these challenges, this article presents NeuroBench, a dual-track, multi-task benchmark framework. NeuroBench addresses the existing neuromorphic benchmark challenges by advancing prior work in three distinct ways. Firstly, the benchmark framework reduces assumptions regarding the specific solution being assessed, encouraging inclusive participation of neuromorphic and non-neuromorphic approaches by utilizing general, task-level benchmarking and hierarchical metric definitions which capture key performance indicators of interest. Secondly, the NeuroBench benchmarks are associated with a common open-source benchmark harness tool which facilitates actionable benchmark implementation and offers structure for further expansion to neuromorphic algorithm frameworks and systems. Finally, NeuroBench establishes an iterative, community-driven initiative designed to evolve over time to ensure representation and relevance to neuromorphic research, analogous to the well-established MLPerf benchmark framework for machine learning [19, 20]. As a whole, NeuroBench intends to align the neuromorphic research community on standard benchmarking, providing a dynamically evolving platform to ensure ongoing relevance and facilitate advancements through workshops, competitions, and a centralized leaderboard.
As Figure 1 shows, the NeuroBench framework involves two tracks to enable agile algorithm and system development. As an emerging technology, neuromorphic hardware has not converged to a single platform which is commercially available, thus a large fraction of neuromorphic research explores algorithmic advancement on conventional systems which may not be optimal for performance. Thus, NeuroBench consists of an algorithm track for hardware-independent evaluation and a system track for fully deployed solutions. The algorithm track defines four novel benchmarks for neuromorphic methods across diverse domains, namely few-shot continual learning, computer vision, motor cortical decoding, and chaotic forecasting, and utilizes complexity metrics to analyze solution costs. Such hardware-independent benchmarking enables algorithmic exploration and prototyping, especially when simulating algorithm execution on non-neuromorphic platforms. Meanwhile, the system track defines standard protocols to measure the real-world speed and efficiency of neuromorphic hardware on benchmarks ranging from standard machine learning tasks to promising fields for neuromorphic systems, such as optimization.
Each NeuroBench track includes defined datasets, metric and measurement methodology, and modular evaluation components to enable flexible development. Promising methods identified from the algorithm track will inform system design by highlighting target algorithms for optimization and relevant system workloads for benchmarking. The system track in turn enables optimization and evaluation of performant implementations, providing feedback to refine algorithmic complexity modeling and analysis. The interplay between the tracks creates a virtuous cycle: algorithm innovations guide system implementation, while system-level insights accelerate further algorithmic progress. This approach allows NeuroBench to advance neuromorphic algorithm-system co-design. Both the algorithm and system track will be extended and co-developed as NeuroBench continues to expand.
In the next few sections, we describe the algorithm track, including general complexity metric definitions, benchmark tasks, and common infrastructure tooling. We apply the framework to report baseline results for each algorithm benchmark, which outline unexplored research opportunities in optimizing algorithmic architectures and training of sparse, stateful models to achieve greater performance and resource efficiency. Then, we show baseline results established in the system track to assess neuromorphic performance across promising application workloads. By outlining both tracks, we provide a roadmap towards standardizing benchmark procedures in both hardware-independent and hardware dependent settings.
## Algorithm Track Benchmark Framework
<details>
<summary>x2.png Details</summary>

### Visual Description
## Diagram: NeuroBench Harness Architecture
### Overview
This diagram illustrates the architectural flow of the "NeuroBench Harness," a system designed to evaluate machine learning models. The diagram is organized into three vertical columns: **Benchmark Inputs** (left), **Benchmark Harness** (center), and **Benchmark Results** (right). It uses a color-coded legend to distinguish between components that are user-defined, user-customizable, and benchmark-defined.
### Components/Axes
The diagram utilizes a legend at the bottom to define the color-coding of the boxes:
* **Green (Solid border):** User-defined
* **Orange (Dashed border):** User-customizable
* **Red (Dotted border):** Benchmark-defined
**Spatial Layout:**
* **Left Column (Benchmark Inputs):** Contains four input blocks.
* **Center Column (Benchmark Harness):** Contains the main processing logic, divided into the "NeuroBenchModel Wrapper" and the "Benchmark Runtime" (which contains a sequence of processing steps).
* **Right Column (Benchmark Results):** Contains two output blocks.
### Detailed Analysis
#### 1. Benchmark Inputs (Left)
* **Model** (Green): Positioned at the top. Feeds into the "NeuroBenchModel Wrapper."
* **Dataset Dataloader** (Red): Positioned below the Model. Feeds into "Load data."
* **Processors Accumulators** (Orange): Positioned below the Dataloader. Feeds into "Apply pre-processing" and "Apply post-processing."
* **Desired metrics** (Orange): Positioned at the bottom. Feeds into "Initialize metric calculations."
#### 2. Benchmark Harness (Center)
This section is divided into two main containers:
* **NeuroBenchModel Wrapper** (Red, top): Acts as an interface. It receives input from the "Model" and sends data to "Calculate static metrics." It also has a dashed-line connection to "Model inference."
* **Benchmark Runtime** (Dark gray container, bottom): Contains a sequential pipeline of operations:
1. **Initialize metric calculations** (Red): Receives input from "Desired metrics."
2. **Load data** (Red): Receives input from "Dataset Dataloader."
3. **Apply pre-processing** (Red): Receives input from "Processors Accumulators."
4. **Model inference** (Red): Receives input from "Apply pre-processing."
5. **Apply post-processing** (Red): Receives input from "Model inference."
* **Metric Calculation Blocks (Right side of Runtime):**
* **Calculate static metrics** (Red): Receives input from "NeuroBenchModel Wrapper."
* **Calculate workload metrics** (Red): Receives input from "Apply post-processing" and "Model inference."
#### 3. Benchmark Results (Right)
* **Static metrics** (Red, top): Receives data from "Calculate static metrics."
* Contains: "Footprint", "Connection sparsity".
* **Workload metrics** (Red, bottom): Receives data from "Calculate workload metrics."
* Contains: "Correctness", "Activation sparsity", "Synaptic operations".
### Key Observations
* **High Standardization:** The vast majority of the "Benchmark Runtime" and "Benchmark Results" components are "Benchmark-defined" (Red), indicating a highly standardized evaluation process.
* **Abstraction Layer:** The "NeuroBenchModel Wrapper" serves as the critical bridge between the user's arbitrary model and the standardized benchmark harness.
* **Data Flow:** The flow is strictly linear for the runtime (Initialize -> Load -> Pre-process -> Inference -> Post-process), while the metric calculations branch off from specific stages of this pipeline.
* **Customization Points:** Users have limited but specific control points: the Model itself, the Dataloader, the hardware/processing configuration ("Processors Accumulators"), and the selection of metrics ("Desired metrics").
### Interpretation
This diagram demonstrates a modular benchmarking framework designed to ensure reproducibility and consistency. By separating the "Benchmark Runtime" from the "Model," the system allows for the evaluation of diverse models using a standardized set of metrics.
The distinction between **Static metrics** and **Workload metrics** is significant:
* **Static metrics** (Footprint, Connection sparsity) appear to be derived from the model's architecture itself, likely analyzed via the "NeuroBenchModel Wrapper" before or independent of execution.
* **Workload metrics** (Correctness, Activation sparsity, Synaptic operations) are derived from the actual execution of the model on data, requiring the full runtime pipeline.
The architecture suggests that the "Benchmark Harness" is designed to be "plug-and-play" for researchers, where they only need to provide the model and data, while the harness handles the complex orchestration of metric calculation and execution.
</details>
Figure 2: An overview of the NeuroBench algorithm track.
The algorithm benchmark track aims to evaluate algorithms in a system-independent manner, separating algorithm performance from specific implementation details. The implementation platform can thus be ill-matched to the particular algorithm benchmark that it executes (e.g., SNN execution via dense matrix multiplication on a GPU), and the algorithm complexity and expected performance can be examined in a theoretical manner, motivating agile prototyping and functional analysis. Furthermore, minimal assumptions are made about the solutions tested, promoting inclusion of diverse algorithmic approaches.
The framework, as illustrated in Figure 2, is composed of inclusively-defined benchmark metrics, datasets and data loaders, and common harness infrastructure, shown in red. The metrics focus on assessing algorithm correctness on specific tasks as well as capturing general metrics that reflect the architectural complexity, computational demands, and storage requirements of the models. The datasets and data loaders specify the details of the tasks used for evaluation and ensure consistency across benchmarks. Finally, the harness infrastructure automates runtime execution and result output for the algorithm benchmark specified by the input interface, which consists of the user’s model and customizable components for data processing and desired metrics, shown in green and orange.
### Algorithm Track Metrics
The algorithm track establishes solution-agnostic primary metrics which are generally relevant to all types of solutions, including artificial and spiking neural networks (ANNs, SNNs). Firstly, there are correctness metrics, which measure the quality of the model predictions on the particular task, such as accuracy, mean average precision (mAP), and mean-squared error (MSE). The correctness metrics are specified per task for each benchmark. Next, there are complexity metrics, which measure the computational demands of the algorithm. In the first iteration of the NeuroBench algorithm track, we assume a digital, time-stepped execution of the algorithm and define the following complexity metrics:
- Footprint – A measure of the memory footprint, in bytes, required to represent a model, which reflects quantization, parameters, and buffering requirements. The metric summarizes (and can be further broken down into) synaptic weight count, weight precision, trainable neuron parameters, data buffers, etc. Zero weights are included, as they are distinguished in the connection sparsity metric.
- Connection Sparsity – For a given model, the connection sparsity is the number of zero weights divided by the total number of weights, accumulated over all layers. 0 refers to no sparsity (fully connected) and 1 refers to full sparsity (no connections). This metric accounts for deliberate pruning and sparse network architectures.
- Activation Sparsity – During execution, the average sparsity of neuron activations over all neurons in all model layers, for all timesteps of all tested samples, where 0 refers to no sparsity (i.e., all neurons are always activated), and 1 refers to the case where all neurons have a zero output.
- Synaptic Operations – Average number of synaptic operations per model execution, based on neuron activations and the associated fanout synapses. This metric is further subdivided into dense, effective multiply-accumulate, and effective accumulate synaptic operations (Dense, Eff_MACs, Eff_ACs). Dense accounts for all zero and nonzero neuron activations and synaptic connections, and reflects the number of operations necessary on hardware that does not support sparsity. Eff_MACs and Eff_ACs only count effective synaptic operations by disregarding zero activations (e.g., produced by the ReLU function in an ANN or no spike in an SNN) and zero connections, thus reflecting operation cost on sparsity-aware hardware. Synaptic operations with non-binary activation are considered multiply-accumulates (MACs), while those with binary activation are considered accumulates (ACs).
Footprint and connection sparsity are classified as static metrics, which can be analytically determined from the model only. Activation sparsity, synaptic operations, and correctness are classified as workload metrics, which are dependent on execution or simulation of the model based on the benchmark data.
In addition to the above complexity metrics, the algorithm track proposes to define Model Execution Rate, corresponding to the rate, in Hz, at which the model’s forward inference pass needs to be executed. For example, if a model is designed to process data from an event camera with a 50 ms input stride, the model execution rate is 20 Hz. The execution rate is a critical feature of the algorithm which provides intuition into the tradeoff between latency and computational footprint of a deployed model, and is reported directly by the solution designer in benchmark results since it neither needs to be calculated nor extracted from the model or its outputs.
The complexity metrics are measured independently of the underlying hardware and therefore do not explicitly correlate with post-deployment latency or energy consumption. However, they provide valuable insight into algorithm performance and resource requirements, enabling high-level comparison and facilitating prototyping. For instance, the execution rate and number of synaptic operations can be taken together to estimate the speed and dynamic power of a model deployed to certain hardware, and the footprint and connection sparsity can be used to proxy hardware resource utilization.
Furthermore, the algorithm track can be extended with solution-specific secondary metrics, which can offer deeper insights by using information specific to particular types of solutions. For example, for algorithms geared towards analog hardware, noise robustness is an important solution-specific metric. In addition, approaches with complex neuron dynamics may warrant measuring the overall complexity of a neuron update (i.e., type and counts of operations necessary to simulate the update), which can be combined with the total number of neuron updates in a model pass to calculate the cost of state updates. Such solution-specific metrics are expected to be community-driven and will be included in future NeuroBench algorithm track releases.
### Algorithm Track Benchmarks
The v1.0 iteration of the NeuroBench algorithm track includes four benchmarks for neuromorphic computing research. The benchmarks were chosen by the NeuroBench community to capture key ongoing challenges for neuromorphic algorithm design. The list of tasks highlights features which are relevant to neuromorphic research interests: few-shot continual learning, object detection utilizing the high dynamic range and temporal resolution of event cameras, sensorimotor decoding based on cortical signals, and low-dimensional predictive modeling useful for prototyping resource-constrained networks that are suitable for small mixed-signal systems. Benchmark tasks are listed below and summarized in Table 1. Detailed specifications of benchmark tasks are provided in the Methods section.
| Task | Dataset | Correctness metric | Task description |
| --- | --- | --- | --- |
| Keyword FSCIL | MSWC [21] | Accuracy | Few-shot, continual learning of keyword classes. |
| Event Camera Object Detection | Prophesee 1MP Automotive [22] | COCO mAP | Detecting automotive objects from event camera video. |
| NHP Motor Prediction | Primate Reaching [23] | R 2 | Predicting fingertip velocity from cortical recordings. |
| Chaotic Function Prediction | Mackey-Glass time series [24] | sMAPE | Autoregressive modeling of chaotic functions. |
Table 1: NeuroBench algorithm track v1.0 benchmarks.
- Keyword Few-Shot Class-Incremental Learning (FSCIL) – Learning new tasks from a small amount of experiences while retaining knowledge of prior tasks is a hallmark of biological intelligence and a long-standing goal of general AI [25]. It is especially a key challenge to endow edge devices with the ability to adapt to their environments and users. This benchmark thus evaluates the capacity of a model to successively incorporate new keywords over multiple sessions (class-incremental), with only a handful of samples from the new classes to train with (few-shot). The FSCIL task is a recently established benchmark in the computer vision domain [26], but it has not yet been adapted to other data modalities. Aligning with a neuromorphic interest in temporal data modalities, this benchmark introduces a FSCIL task with streaming audio data using the large Multilingual Spoken Word Corpus (MSWC) [21] keyword classification dataset. The task is designed to be approached in two phases: pre-training and incremental learning. First, for pre-training, a set of 100 words spanning 5 base languages (English, German, Catalan, French, Kinyarwanda) with 500 training samples each are made available to train an initial model. Next, for incremental learning, the model undergoes 10 successive sessions to learn words from 10 new languages (Persian, Spanish, Russian, Welsh, Italian, Basque, Polish, Esparanto, Portuguese, Dutch) in a few-shot learning scenario. Each incremental session adds 10 words of the corresponding session language with only 5 training samples available per word. After each session, the model is tested in classification accuracy on all prior learned classes, including the 100 base pre-training classes and the few-shot-learned classes, therefore evaluating the FSCIL solution on its ability to learn new classes while retaining knowledge about the previously learned ones. Each session learns a new language, for a total knowledge base of 200 keywords by the end of the benchmark.
- Event Camera Object Detection – Object detection is a widely-used computer vision task with applications in robotics, autonomous driving, and surveillance. Such scenarios at the edge may require high energy efficiency and real-time performance, which can be achieved via event-based vision sensors [27]. The event camera object detection benchmark uses the Prophesee 1 Megapixel automotive detection dataset [22], a large labeled object detection dataset with over 15 hours of event camera video from the front of a car driving in various scenarios. Predetermined training, validation, and testing splits include $11.2 h$ , $2.2 h$ , and $2.2 h$ of recording, respectively. Pedestrian, two-wheeler, and car object classes are used in evaluation, and correctness is measured using COCO mean average precision (mAP) [28].
- Non-human Primate (NHP) Motor Prediction – Studying models which can accurately replicate features of biological computation presents opportunities in understanding sensorimotor behavior and developing closed-loop methods for future robotic agents. It also is foundational to the development of wearable or implantable neuro-prosthetic devices that can accurately generate motor activity from neural or muscle signals. This benchmark utilizes a dataset consisting of multi-channel recordings from the sensorimotor cortex of two non-human primates (NHP Indy and NHP Loco) during reaching movements, along with corresponding fingertip motion of the reach [23]. Six total sessions are included from the dataset, for a total of 8712 seconds of data. The task is to train a model to predict the two-dimensional components of finger velocity using recent neural data. The sessions are treated independently (i.e., models are trained separately for each session), and the data is split to allow the first 75% for training and validation and the last 25% for evaluation. Correctness of the predictions is evaluated by the coefficient of determination ( $R^2$ ) score against the true finger velocity targets, averaged over all six sessions.
- Chaotic Function Prediction – The real-world data benchmarks presented thus far are high-dimensional and can require large networks to achieve high accuracy, raising challenges for solution types with limited I/O support and network capacity, such as mixed-signal edge prototype solutions. To address this, we include a synthetic benchmark based on prediction of one-dimensional Mackey-Glass time series [24], which can be effectively tackled by smaller networks. Mackey-Glass has been widely adopted as a benchmark for evaluating temporal predictors, including neuromorphic models [29, 30, 31]. The task involves prediction of the next timestep value $f(t+Δ t)$ given the current timestep value $f(t)$ . The model is trained and validated using the first half of the time series, during which the ground truth state $f(t)$ are supplied to the model to predict the next timestep $f^\prime(t+Δ t)$ . During the evaluation, the model uses its prior prediction $f^\prime(t)$ to generate each next value $f^\prime(t+Δ t)$ , autoregressively forecasting the second half of the time series. Correctness is measured using symmetric mean absolute percentage error (sMAPE) of the generated time series against the target time series, a standard metric in forecasting [32]. The benchmark includes a set of 14 Mackey-Glass time series, which vary by the equation parameter $τ$ , the delay constant. Lyapunov time $(L)$ , the expected predictability timescale for chaos [33], is used as the time unit for each time series. The total length of each series is 20 Lyapunov times, and 75 points are sampled per Lyapunov time ( $Δ t$ = $L/75$ ).
### Algorithm Track Benchmark Harness
The NeuroBench algorithm benchmarks are wrapped in a harness which standardizes the benchmark interfaces. The harness provides benchmark users with a consistent framework for loading data, processing data and model outputs, and calculating and reporting metrics, thereby ensuring fair and standard comparisons of the results. It is built with straightforward interfaces which are designed to be extended with new frameworks, algorithms, and tasks. The benchmark harness is open-source for use and development (https://github.com/NeuroBench/neurobench).
The components of the algorithm benchmark harness are summarized in Figure 2. Datasets are loaded in a common format and pass through Processors to be pre-processed. The Model generates predictions based on the processed data, and Accumulators post-process the predictions, for instance to accumulate spikes and transform to labels. Static metrics of algorithm footprint and connection sparsity are calculated via model analysis, while metrics of correctness, activation sparsity, and synaptic operations are calculated using predictions and model execution traces. For benchmark users, task evaluation simply involves utilizing the existing dataloaders, processors, and metrics within the harness and wrapping their own code to fit the standard interfaces.
Currently, the harness and all baseline models are built using PyTorch [34] or frameworks based on it, such as snnTorch [18] and SpikingJelly [35]. Due to its modular structure and simple interfaces, the harness can grow to be compatible with further neuromorphic tools such as Lava [36] and Fugu [37]. Furthermore, it also supports the extension of data and metric pipelines in order to implement additional benchmark tasks. Widely validated benchmarks in keyword [38] and gesture classification [12], which are foundational in neuromorphic and conventional research [39, 13, 40, 41], have been incorporated into the harness to complement the novel tasks in the NeuroBench v1.0 suite. Any novel or existing benchmarks can make use of the harness infrastructure for open reproducibility, and also to garner interest in the community towards long-term task support and appearance in NeuroBench-affiliated leaderboards and challenge events.
#### Algorithm Track Limitations and Further Extensions
Before diving into the baseline results, it is worth discussing several possible improvements to the NeuroBench algorithm track framework in its current form. Specifically, the initial iteration of metrics is restricted to the assumption of digital, time-stepped algorithm execution. While complexity analysis of such prototypes can serve as an intermediate step for solutions intended for analog or continuous time deployment, the metric measurements are not yet defined for those execution settings. Informed by further benchmark implementations, future versions of NeuroBench will extend inclusiveness by expanding measurement protocols to include such algorithms.
Furthermore, the synaptic operations metric, intended to capture model computation cost, currently does not account for neuron updates. The dynamics of neuron models, including mechanisms like leakage and reset, can vary heavily in complexity. However, counting the number and type of operations from neuron updates, as well as estimating their overall costs, depends on the specific arithmetic or circuit implementation. Thus, they are not accounted for in the broader algorithmic complexity metrics. The algorithmic metric framework can be extended with solution-specific metrics that assume a particular implementation platform to estimate neuron update costs, which have been previously defined [42]. These estimates can then be combined with the total number of neuron updates per model computation to measure overall network operation complexity during evaluation.
Data pre- and post-processing can also amount to significant costs not yet captured in the NeuroBench algorithm track metrics. Such costs are, however, captured in the deployed metrics of the system track, which accounts for data processing hardware as part of the overall system during performance and efficiency measurements. Data processing metrics will be added as a separate complexity category for the algorithm track benchmark in the future.
The v1.0 algorithm track benchmark suite is also intended to expand in the future. This could include covering further data modalities such as inertial measurement unit (IMU) sensing [43] and extending to closed-loop sensorimotor tasks to demonstrate embodied intelligence. As with the initial benchmarks, further tasks will undergo approval and development by the open NeuroBench community before being included in a future versioned benchmark suite.
### Algorithm Track Baseline Results
In our first iteration of the algorithmic track, we report baseline algorithm performance on each benchmark using various model architectures, including artificial neural networks commonly used in deep learning, spiking neural networks, and reservoir networks. We evaluate each benchmark with two substantially different algorithm baselines. From these evaluations, we extract baseline comparisons, identify trends, and uncover motivations for future research. Except for the event camera object detection task, each benchmark utilizes a novel data split, and all tasks use novel metric measurement. The presented baselines are a snapshot of the solution search space and will be starting points for leaderboards, thereby calling for further research to push the state of the art for each task. Detailed specifications of each of the baselines can be found in the Methods section.
#### Keyword FSCIL
Baseline Accuracy (Base / Session Avg) Footprint (bytes) Model Exec. Rate (Hz) Connection Sparsity Activation Sparsity SynOps (per model exec.) Dense Eff_MACs Eff_ACs M5 ANN (97.09% / 89.27%) $6.03× 10^6$ 1 0.0 0.783 $2.59× 10^7$ $7.85× 10^6$ 0 SNN (93.48% / 75.27%) $1.36× 10^7$ 200 0.0 0.916 $3.39× 10^6$ 0 $3.65× 10^5$
Table 2: Baseline results for the keyword few-shot class-incremental learning task. Base accuracy refers to accuracy on the 100 base classes after pre-training while session average accuracy is the average accuracy over all sessions for the corresponding prototypical baseline. The detailed accuracy per session for the different baselines are shown in Figure 3.
The keyword FSCIL task has an ANN and SNN baseline, using different model architectures:
- M5 ANN – The ANN baseline uses a tuned version of the M5 deep convolutional network architecture [44], with samples pre-processed into Mel-frequency cepstral coefficients (MFCC). The network contains four successive convolution-normalization-pooling layers, followed by a readout fully-connected layer. Each model execution (forward pass) uses the data from the full pre-processed sample, and convolution kernels are applied over the temporal dimension of the samples. This is reported as a 1 Hz model execution rate.
- SNN – The SNN baseline uses a recurrent SNN with adaptive leaky integrate-and-fire (LIF) neurons and heterogeneous time constants [45]. The SNN consists of two recurrent adaptive LIF layers and one linear output layer. Audio samples are pre-processed to binary spike trains using Speech2Spikes [46], which relies on a Mel Spectrogram with the same parameters as the MFCC of the ANN baseline. Each input timestep to the model represents 5 ms of audio data, thus the model has a 200 Hz model execution rate. Output neuron activations are summed over time to produce the word class prediction.
After pre-training using standard batched training, the ANN and SNN baseline networks reach high accuracies on the base classes of 97.09% and 93.48%, respectively. As reported by the model execution rate metric, the SNN baseline computes each sample over 200 passes, using an order of magnitude fewer effective AC synaptic operations compared to the ANN baseline’s effective MACs per model execution. Considering both the model execution rate and synaptic operation metrics, the number of aggregated ACs over the length of the sample ( $200*3.65× 10^5=7.30× 10^7$ ) exceeds the Dense and effective MAC operations necessary for the ANN baseline, which spatially flattens the sample and processes it in one model execution. However, outside of the static-length keyword classification scenario, the low-cost per-execution temporal processing of SNNs can enable efficient, always-on, high-frequency prediction capabilities in deployed continuous audio recognition scenarios.
<details>
<summary>extracted/6132287/figures/FSCIL_proto.png Details</summary>

### Visual Description
## Line Charts: Performance Comparison of Prototypical and Frozen Models
### Overview
This image displays two line charts comparing the performance of four different machine learning model configurations ("Prototypical M5 ANN", "Prototypical SNN", "Frozen M5 ANN", and "Frozen SNN") across incremental learning sessions. The left chart illustrates performance on "All Classes," while the right chart illustrates performance on "New Classes." Both charts utilize shaded regions to represent variance or uncertainty around the mean performance values.
### Components/Axes
**Common Elements:**
* **Y-Axis:** Labeled "Test Accuracy (%)", ranging from 40 to 100.
* **Grid:** Both charts feature a light gray grid for easier value estimation.
**Left Chart: "All Classes Performance"**
* **X-Axis:** "Incremental Sessions", ranging from 0 to 10.
* **Legend (Bottom-Left):**
* **Prototypical M5 ANN:** Red solid line with circular markers.
* **Prototypical SNN:** Blue solid line with circular markers.
* **Frozen M5 ANN:** Red dashed line.
* **Frozen SNN:** Blue dashed line.
**Right Chart: "New Classes Performance"**
* **X-Axis:** "Incremental Sessions", ranging from 1 to 10.
* **Legend (Bottom-Left):**
* **Prototypical M5 ANN:** Red solid line with circular markers.
* **Prototypical SNN:** Blue solid line with circular markers.
---
### Detailed Analysis
#### Left Chart: All Classes Performance
* **Trend Verification:** The "Prototypical" models (solid lines) show a gradual, linear-like decline. The "Frozen" models (dashed lines) show a steep, exponential-like decay.
* **Prototypical M5 ANN (Red Solid):** Starts at approximately 95% at session 0. It declines steadily to approximately 85% by session 10.
* **Prototypical SNN (Blue Solid):** Starts at approximately 84% at session 0. It declines steadily to approximately 69% by session 10.
* **Frozen M5 ANN (Red Dashed):** Starts at approximately 97% at session 0. It drops sharply, reaching ~75% by session 4, and continues to fall to approximately 48% by session 10.
* **Frozen SNN (Blue Dashed):** Starts at approximately 93% at session 0. It drops sharply, reaching ~67% by session 4, and continues to fall to approximately 47% by session 10.
#### Right Chart: New Classes Performance
* **Trend Verification:** Both models show a very slight, gradual decline across the sessions.
* **Prototypical M5 ANN (Red Solid):** Starts at approximately 82% at session 1. It maintains a high level of performance, ending at approximately 78% by session 10.
* **Prototypical SNN (Blue Solid):** Starts at approximately 59% at session 1. It maintains a lower, but stable performance, ending at approximately 56% by session 10.
---
### Key Observations
* **Catastrophic Forgetting:** The "Frozen" models (left chart) exhibit severe performance degradation, characteristic of catastrophic forgetting in incremental learning, where the model loses the ability to classify previous classes as new ones are added.
* **Stability:** The "Prototypical" models demonstrate significantly higher stability in both "All Classes" and "New Classes" scenarios compared to the "Frozen" counterparts.
* **Model Superiority:** In all observed scenarios, the "M5 ANN" architecture consistently outperforms the "SNN" architecture.
* **Variance:** The shaded regions indicate that the "Prototypical M5 ANN" has a tighter variance (more consistent performance) in the "All Classes" chart compared to the "Prototypical SNN." In the "New Classes" chart, the "Prototypical M5 ANN" shows a wider variance (shaded red area) compared to the "Prototypical SNN" (shaded blue area).
### Interpretation
The data strongly suggests that the "Prototypical" approach is superior to the "Frozen" approach for incremental learning tasks. The "Frozen" models, which likely utilize a fixed feature extractor, fail to adapt to the distribution shifts inherent in adding new classes, leading to a rapid collapse in accuracy.
Conversely, the "Prototypical" models maintain a much higher degree of knowledge retention (as seen in the "All Classes" chart) and consistent performance on new data (as seen in the "New Classes" chart). The M5 ANN architecture appears to be more robust or better suited for this specific task than the SNN architecture, consistently yielding higher accuracy across all metrics. This visualization effectively demonstrates the trade-off between model rigidity ("Frozen") and adaptability ("Prototypical").
</details>
Figure 3: Test accuracy per session on the keyword FSCIL task for prototypical and frozen baselines, with the accuracy on both base classes and incrementally-learned classes (left), and accuracy on all incrementally-learned classes only (right). Incremental session 0 refers to the accuracy on base classes after pre-training only. Shaded area represents 5 th and 95 th percentile on 100 runs. Frozen baselines with no adaptation do not learn incremental classes and thus have a fixed 0% accuracy for New Classes Performance.
We present two approaches for the incremental stage for both the ANN and SNN baselines. The frozen models are locked after pre-training on base classes and have 0% accuracy on all new incremental classes, providing a reference for models with no learning or catastrophic forgetting of prior classes. The prototypical models employ a prototypical network [47] for incremental learning, which is a feature-based clustering approach that can be implemented with a simple linear readout layer on top of the pre-trained network backbone. Prototypical weights and biases of prior and incremental classes are directly defined based on the average features of the corresponding class and directly substitute pre-trained readout layer parameters. The complexity results in Table 2 thus empirically apply to both the frozen and prototypical models.
The test accuracy for the baseline models over all sessions, as well as the test accuracy on only the new incrementally-learned classes, are shown in Figure 3. Using prototypical networks, the ANN model reaches 89.27% accuracy on average over all sessions, demonstrating significant greater performance of 21.41 accuracy points with respect to the frozen model. The accuracy on new classes, averaged over all incremental sessions, is 79.61%. The SNN prototypical baseline, on the other hand, reaches 75.27% accuracy on average over all sessions, surpassing the frozen SNN performance by 9.97 accuracy points, with an average accuracy on new classes over all sessions of 57.23%.
The accuracy loss over the incremental sessions is similar between the ANN and SNN prototypical baselines. However, the lower overall accuracy of the SNN is largely due to the conversion from the original backpropagation-trained readout classifier, which is used in the frozen baseline, to the prototype readout classifier. On the base classes (session 0 in Figure 3), the ANN sees a drop of 2.37% between the frozen and prototypical baselines, while the SNN has a larger drop of 9.17%. The larger drop indicates that our particular SNN baseline has a less general feature extraction than the ANN. This may be due to the challenges of backpropagation through time for online temporal inference to learn to extract long-term temporal keyword features with the chosen spiking recurrent model. Additionally, the Speech2Spikes [46] pre-processing algorithm converting audio to spikes may also cause information loss. Overall, the keyword FSCIL benchmark presents opportunities for further research in learning methods, preprocessing, and model architectures for continual learning of temporal data.
#### Event Camera Object Detection
Baseline mAP Footprint (bytes) Model Exec. Rate (Hz) Connection Sparsity Activation Sparsity SynOps (per model exec.) Dense Eff_MACs Eff_ACs RED ANN 0.429 $9.13× 10^7$ 20 0.0 0.634 $2.84× 10^11$ $2.48× 10^11$ 0 Hybrid 0.271 $1.21× 10^7$ 20 0.0 0.613 $9.85× 10^10$ $3.76× 10^10$ $5.60× 10^8$
Table 3: Baseline results for the event camera object detection task.
The event camera object detection task reports a prior baseline, the RED ANN, and a novel conversion of the architecture to a hybrid ANN-SNN model:
- RED ANN – The RED architecture [22] consists of blocks of feed-forward squeeze-and-excite [48] convolutional layers followed by blocks of recurrent convolution-LSTM (ConvLSTM [49]) layers. A single-shot detection (SSD [50]) head is used to predict the location and class of the bounding box based on multi-scale outputs from the recurrent layers. Raw event data is binned into 50 ms and pre-processed into time surfaces.
- Hybrid – The hybrid ANN-SNN architecture adopts feedforward LIF spiking neural layers to replace the ConvLSTM layers in RED, and shares the same feed-forward convolutional blocks as the RED. It uses the same input encoding method and SSD head as the RED model.
Results for the two networks can be found in Table 3. The RED ANN represents the current state-of-the-art correctness on the benchmark, at 0.429 mAP. The Hybrid network is a smaller network, reflected by the footprint and synaptic operations metrics measuring an order of magnitude smaller than for the RED ANN. The smaller size comes at the expense of lower correctness of 0.271 mAP.
For the RED ANN, the activation sparsity metric (0.634) represents zero activations by the ReLU function for each neuron. From this, one may expect that the number of effective operations (operations with a nonzero activation and nonzero weight) would be around 35% of dense operations, however the actual ratio is 87%. This is due to the presence of normalization layers applied to activations before synaptic weight multiplication. Furthermore, neurons with lower activation frequency in the network tend to have a smaller fanout than neurons with high activation frequency. Thus, while activation sparsity alone can provide a proxy for the cost of the network, architectural characteristics may impede actual computation reduction, and the synaptic operations must be considered in tandem.
The Hybrid network demonstrates a significant reduction in total effective operations against dense operations, outlining significant gains if deployed on specialized sparsity-aware hardware. However, for the particular network, the number of effective ACs, generated by the spiking neuron components, is two orders of magnitude smaller than the number of effective MACs within the ANN components. Such a hybrid network may not warrant specialized accumulation units, and the baseline motivates further research in hybrid networks with a larger proportion of spiking neuron activity compared to artificial neuron activity.
#### NHP Motor Prediction
Baseline $R^2$ Footprint (bytes) Model Exec. Rate (Hz) Connection Sparsity Activation Sparsity SynOps (per model exec.) Dense Eff_MACs Eff_ACs ANN 0.593 20824 250 0.0 0.683 4704 3836 0 0.558 33496 250 0.0 0.668 7776 6103 0 SNN 0.593 19648 250 0.0 0.997 4900 0 276 0.568 38848 250 0.0 0.999 9700 0 551
Table 4: Baseline results for the NHP motor prediction task, for NHP Indy (96-channel data, top), and NHP Loco (192-channel data, bottom).
Small fully-connected, feedforward networks were developed for the NHP motor prediction baselines:
- ANN – In the ANN baseline, the cortical activity from the 50 most recent data samples is buffered to be used as network input. The network has two hidden layers and $2$ final outputs predicting $X$ and $Y$ velocities, with a fully-connected topology of $N_ch$ -32-48-2, where $N_ch$ refers to the channels of cortical data (96 for NHP Indy, and 192 for NHP Loco). Batch normalization is applied after each hidden layer.
- SNN – The SNN uses the data samples directly as input to the network, without buffering. It has a hidden layer of 50 LIF neurons, for a fully connected topology of $N_ch$ -50-2 LIF neurons. The output neurons do not have a reset mechanism, and the membrane potential is directly read to produce the output velocities.
Table 4 shows the results for the ANN and SNN baselines, averaged between sessions from each NHP (Indy and Loco). The ANN and SNN are similar in footprint size and number of dense operations per model forward pass, and also reach comparable prediction quality based on $R^2$ score. Each model is small in footprint and operation count, demonstrating that this task can be solved by shallow edge networks, validating prior studies [51].
Between the baselines, the SNN realizes similar correctness at significantly reduced complexity compared to the ANN. Extremely high activation sparsity in the SNN (0.998) directly translates to low effective accumulate operations, demonstrating the adequacy of stateful, binary-activation neuron models for sparse regression tasks. Meanwhile, similarly to the RED ANN in the event camera object detection task, activation sparsity in the ANN baseline does not translate to effective operation efficiency, as batch normalization is applied to activations before multiplication with synaptic weights.
<details>
<summary>extracted/6132287/figures/Memory_vs_Accuracy.png Details</summary>

### Visual Description
## Scatter Plot: Model Footprint vs. $R^2$ Performance
### Overview
This image is a scatter plot comparing the memory "Footprint (bytes)" on the Y-axis against the "$R^2$" (coefficient of determination) on the X-axis for four distinct model types. The plot utilizes a logarithmic scale for the Y-axis. There is an unexplained distinction between hollow and filled markers for each category, which is not defined in the provided legend.
### Components/Axes
* **Y-Axis:** Labeled "Footprint (bytes)". The scale is logarithmic, with a visible tick mark at $10^5$.
* **X-Axis:** Labeled "$R^2$". The scale ranges from 0.56 to 0.64.
* **Legend (Located in the top-left quadrant):**
* **Square:** ANN
* **Star:** SNN
* **Diamond:** ANN_Flat
* **Plus:** SNN_Flat
### Detailed Analysis
The plot contains eight data points in total, consisting of four categories, each represented by both a hollow and a filled marker.
| Category | Marker Type | Approx. $R^2$ (X) | Approx. Footprint (Y) |
| :--- | :--- | :--- | :--- |
| **ANN** | Hollow Square | 0.56 | $3 \times 10^4$ |
| **ANN** | Filled Square | 0.595 | $2 \times 10^4$ |
| **SNN** | Hollow Star | 0.575 | $3.5 \times 10^4$ |
| **SNN** | Filled Star | 0.59 | $1.8 \times 10^4$ |
| **ANN_Flat** | Hollow Diamond | 0.59 | $4 \times 10^5$ |
| **ANN_Flat** | Filled Diamond | 0.64 | $9 \times 10^4$ |
| **SNN_Flat** | Hollow Plus | 0.625 | $4 \times 10^4$ |
| **SNN_Flat** | Filled Plus | 0.64 | $2.5 \times 10^4$ |
*Note: Y-axis values are estimated based on the logarithmic scale relative to the $10^5$ marker.*
### Key Observations
* **Performance Clustering:** The "Flat" variants (ANN_Flat and SNN_Flat) consistently achieve higher $R^2$ values (ranging from ~0.59 to 0.64) compared to the standard ANN and SNN models (ranging from ~0.56 to 0.595).
* **Outlier:** The hollow diamond (ANN_Flat) is a significant outlier in terms of memory usage, with a footprint significantly higher than all other data points ($>10^5$ bytes).
* **Efficiency:** The filled SNN_Flat marker represents a highly efficient model, achieving the highest $R^2$ ($\approx 0.64$) while maintaining a relatively low footprint ($\approx 2.5 \times 10^4$ bytes).
* **Marker Discrepancy:** The legend does not define the difference between hollow and filled markers. However, across all categories, the filled markers generally show a lower footprint than their hollow counterparts (with the exception of the ANN_Flat, where the filled marker is lower, but the hollow marker is an extreme outlier).
### Interpretation
This data demonstrates a trade-off analysis between model complexity (footprint) and predictive accuracy ($R^2$).
1. **Optimization Impact:** The "Flat" architecture appears to be a superior design choice for maximizing $R^2$, as both ANN_Flat and SNN_Flat occupy the right side of the X-axis.
2. **Efficiency:** The SNN_Flat model appears to be the most desirable configuration. It achieves the same peak $R^2$ as the ANN_Flat (filled) but does so with a significantly smaller memory footprint (roughly 1/4th the size of the filled ANN_Flat).
3. **Experimental Conditions:** The presence of both hollow and filled markers suggests two different experimental conditions (e.g., "Baseline" vs. "Optimized" or "Training" vs. "Inference"). The consistent shift in the filled markers suggests that the second condition generally results in a lower memory footprint across all model types.
</details>
<details>
<summary>extracted/6132287/figures/Computes_vs_Accuracy.png Details</summary>

### Visual Description
## Scatter Plot: Effective Synaptic Operations vs. $R^2$
### Overview
This image is a scatter plot visualizing the trade-off between model performance (measured by $R^2$) and computational cost (measured by "Effective Synaptic Operations"). The plot compares four distinct model architectures: ANN, SNN, ANN_Flat, and SNN_Flat. The Y-axis uses a logarithmic scale.
### Components/Axes
* **X-Axis:** Labeled "$R^2$". The scale ranges from 0.56 to 0.64.
* **Y-Axis:** Labeled "Effective Synaptic Operations". The scale is logarithmic, spanning from $10^3$ to $10^4$ (and extending beyond $10^4$).
* **Legend:** Located in the bottom-right quadrant. It defines four marker types:
* **Square (filled):** ANN
* **Star (filled):** SNN
* **Diamond (filled):** ANN_Flat
* **Plus (filled):** SNN_Flat
* **Note on Markers:** The plot contains both filled and hollow versions of these shapes. The legend only provides keys for the filled versions; the distinction between hollow and filled markers is not explicitly defined in the image.
### Detailed Analysis
The following data points are estimated based on the visual positioning relative to the axes:
| Model Type | Marker Style | Approx. $R^2$ (X) | Approx. Synaptic Ops (Y) |
| :--- | :--- | :--- | :--- |
| **ANN** | Hollow Square | 0.56 | ~6,000 |
| **ANN** | Filled Square | 0.59 | ~4,000 |
| **SNN** | Hollow Star | 0.575 | ~600 |
| **SNN** | Filled Star | 0.59 | ~300 |
| **ANN_Flat** | Hollow Diamond | 0.59 | ~15,000 |
| **ANN_Flat** | Filled Diamond | 0.64 | ~8,000 |
| **SNN_Flat** | Hollow Plus | 0.625 | ~40,000 |
| **SNN_Flat** | Filled Plus | 0.64 | ~25,000 |
### Key Observations
* **Computational Cost Hierarchy:** The "Flat" variants (Diamond and Plus) consistently occupy the upper region of the chart, indicating significantly higher "Effective Synaptic Operations" compared to the standard ANN and SNN models.
* **Efficiency Gap:** SNN (Star) models are positioned at the lowest Y-values, suggesting they are the most computationally efficient architectures among those plotted.
* **Performance vs. Cost:** There is a general trend where higher $R^2$ values (moving right on the X-axis) correlate with higher computational costs (moving up on the Y-axis), particularly when comparing the "Flat" models to the standard models.
* **Ambiguous Variable:** The distinction between hollow and filled markers is consistent across all four categories (each has one hollow and one filled marker). This suggests a secondary variable (e.g., training vs. testing data, or different hyperparameter configurations) that is not labeled in the legend.
### Interpretation
This chart demonstrates a classic Pareto-style trade-off analysis in neural network architecture design.
1. **Efficiency vs. Accuracy:** The data suggests that while "Flat" architectures (ANN_Flat, SNN_Flat) achieve higher $R^2$ scores (approaching 0.64), they do so at a substantial computational cost, requiring an order of magnitude more synaptic operations than their standard counterparts.
2. **SNN Advantage:** The SNN models appear to be the most efficient, operating well below the $10^3$ threshold, though they also appear to have lower $R^2$ scores compared to the "Flat" variants.
3. **Missing Context:** The presence of both hollow and filled markers without a legend definition is a significant omission. Given the consistent pairing (one hollow, one filled per category), it is highly probable that these represent two different experimental conditions (e.g., "Training" vs. "Inference" or "Model A" vs. "Model B"). Without this key, the full relationship between the data points cannot be definitively determined.
</details>
Figure 4: Footprint and effective synaptic operations vs $R^2$ , for four task baselines. Each model has two points: the solid marker represents NHP Indy, and the hollow marker represents NHP Loco.
We conduct further exploration for increasing task accuracies with more complex ANN and SNN models: ANN_Flat and SNN_Flat. For these networks, 50 data samples of buffered input are split into $n_p=7$ accumulated bins. For ANN_Flat, the 7 bins are spatially flattened as input to the network, so its topology is ( $7× N_ch$ )-32-48-2. SNN_Flat uses the $N_ch$ -32-48-2 topology, and the 7 bins are temporally flattened as input, presented to the network as separate input timesteps. Each prediction still uses the membrane potential of the output neurons after input timesteps, and the network is reset for each prediction. Layer normalization is also applied on the SNN_Flat inputs.
Figure 4 shows plots of complexity and predictive quality of all four baseline networks. Both flattened networks demonstrate significantly greater $R^2$ performance than the other two networks. However, the larger input dimension of the ANN_Flat network is reflected in its greater footprint, and the increased model timesteps and layer normalization sharply increase the effective operations of SNN_Flat by two orders of magnitude compared to the simpler SNN. Thus, while input flattening and normalization increase the quality of model predictions for ANNs and SNNs, each comes with a significant complexity trade-off.
#### Chaotic Function Prediction
Baseline sMAPE Footprint (bytes) Model Exec. Rate (Hz) Connection Sparsity Activation Sparsity SynOps (per model exec.) Dense Eff_MACs Eff_ACs ESN 14.79 $2.81× 10^5$ - 0.876 0.0 $3.52× 10^4$ $4.37× 10^3$ 0 LSTM 13.37 $4.90× 10^5$ - 0.0 0.530 $6.03× 10^4$ $6.03× 10^4$ 0
Table 5: Baseline results for the chaotic function prediction task. Execution rate is not reported as the data is a synthetic time series, with no real-time correlation.
The chaotic function prediction task has two recurrent ANN baselines, which feature distinct network architectures:
- Long short-term memory (LSTM) – LSTMs are a class of recurrent ANN architectures [52], utilizing multiple gates for selective retention or omission of past information. The LSTM baseline consists of a single LSTM with a hidden state of 100 neurons, followed by a feed-forward layer to produce single-dimension output predictions. In addition, the LSTM baseline utilizes explicit memory by buffering 50 previous datapoints, spatially flattening them into 50 input channels.
- Echo state network (ESN) – ESNs are randomized recurrent ANNs that belong to a class of algorithms known collectively as reservoir computing [53], featuring more biologically-inspired principles than LSTMs despite not being spiking networks. Standard ESNs have only one hidden layer (the reservoir), where synaptic connections projecting input data to the hidden layer and recurrent synaptic connections within the hidden layer are chosen randomly and stay fixed during the training. The model architecture for the ESN baseline has two neurons in the input layer, which projects the Mackey-Glass function input and additional constant bias input into a hidden layer of 186 neurons. Within the hidden layer, the probability of recurrent connections is set to 0.11.
The LSTM and ESN models were evaluated on a Mackey-Glass time series with $τ=17$ . The model is evaluated over 30 instantiations of the system; in each instance the start point is shifted forward by half of the Lyapunov time. The model is re-initialized and re-trained on each instance, and the results are averaged over all 30 instances.
The ESN model is architecturally unique compared to the other ANN and SNN baselines. The connection sparsity metric ( $0.876$ ) reflects the high number of zero-weight connections across its reservoir hidden layer. Due to this sparsity, hardware with support for sparse synaptic representation by ignoring zero weights would require less memory to represent the network, thus decreasing the deployed footprint of the model. The high connection sparsity of the ESN leads to significant reduction in synaptic operations - the ESN uses an order of magnitude fewer effective operations ( $4.37× 10^3)$ than the LSTM ( $6.03× 10^4$ ), while achieving comparable sMAPE. The activation sparsity of the ESN is 0 due to neurons using $\tanh(·)$ , rather than ReLU activations.
<details>
<summary>x3.png Details</summary>

### Visual Description
## Scatter Plot: Performance Comparison of ESN vs. LSTM Models
### Overview
This image is a scatter plot comparing the performance of two machine learning models, **ESN** (Echo State Network) and **LSTM** (Long Short-Term Memory), across a range of values for the parameter **$\tau$** (tau). The performance is measured using **sMAPE** (Symmetric Mean Absolute Percentage Error), where lower values indicate better model performance.
### Components/Axes
* **Y-Axis (Vertical):** Labeled "sMAPE". The scale ranges from 0 to 200, with major grid lines at intervals of 50.
* **X-Axis (Horizontal):** Labeled "$\tau$". The scale ranges from 17 to 30 (implied by the data points), with major grid lines labeled at 18, 20, 22, 24, 26, 28, and 30.
* **Legend:** Located at the top center.
* **Blue Circle:** Represents the **ESN** model.
* **Red Circle:** Represents the **LSTM** model.
### Detailed Analysis
The data points are plotted at integer intervals for $\tau$ from 17 to 30. Below are the approximate values for sMAPE for each model:
| $\tau$ | ESN (Blue) sMAPE | LSTM (Red) sMAPE |
| :--- | :--- | :--- |
| **17** | ~15 | ~15 |
| **18** | ~18 | ~18 |
| **19** | ~68 | ~26 |
| **20** | ~54 | ~26 |
| **21** | ~38 | ~28 |
| **22** | ~52 | ~30 |
| **23** | ~93 | ~33 |
| **24** | ~128 | ~28 |
| **25** | ~112 | ~30 |
| **26** | ~118 | ~32 |
| **27** | ~158 | ~35 |
| **28** | ~162 | ~33 |
| **29** | ~143 | ~30 |
| **30** | ~145 | ~34 |
**Trend Verification:**
* **LSTM (Red):** The trend is relatively flat and stable. It maintains a low sMAPE (between ~15 and ~35) across the entire range of $\tau$, showing high robustness.
* **ESN (Blue):** The trend is highly volatile and shows a general upward slope. While it starts at parity with LSTM at $\tau=17-18$, it diverges sharply at $\tau=19$ and continues to exhibit significant fluctuations and an overall increase in error as $\tau$ increases.
### Key Observations
* **Performance Parity:** At low values of $\tau$ (17 and 18), both models perform almost identically with very low sMAPE values.
* **Divergence:** Starting at $\tau=19$, the ESN model's error increases significantly, while the LSTM model remains stable.
* **Volatility:** The ESN model exhibits high variance in its error rate (e.g., dropping from ~128 at $\tau=24$ to ~112 at $\tau=25$, then rising again), whereas the LSTM model shows a very consistent, tight clustering of error values.
* **Outliers/Anomalies:** The ESN model shows a distinct "spike" in error at $\tau=24$ and $\tau=27-28$, suggesting instability in the model's predictive capability at these specific parameter settings.
### Interpretation
The data demonstrates that the **LSTM model is significantly more robust and accurate** than the ESN model for this specific task as the parameter $\tau$ increases.
* **Model Suitability:** The ESN model appears to be highly sensitive to the $\tau$ parameter. Its performance degrades rapidly as $\tau$ increases, suggesting it may struggle to capture long-term dependencies or patterns in the data as the embedding dimension or time lag grows.
* **Stability:** The LSTM model's consistent performance suggests it is better suited for this dataset, as it maintains a low error rate regardless of the $\tau$ value.
* **Peircean Investigative Note:** The divergence at $\tau=19$ acts as a "critical point." Before this point, the models are indistinguishable. After this point, the ESN's failure to maintain low error suggests a fundamental limitation in the ESN architecture's ability to handle the complexity introduced by higher $\tau$ values, whereas the LSTM architecture successfully generalizes across the tested range.
</details>
Figure 5: ESN and LSTM models evaluated on varying Mackey-Glass time series using a constant set of hyperparameters.
Furthermore, we show the generalization and robustness capabilities of the particular ESN and LSTM models by applying them, with fixed hyperparameter sets, to other Mackey-Glass time series. Figure 5 shows the sMAPE score of the models over varied time series with the $τ$ Mackey-Glass parameter varying between 17 and 30. The models were trained independently for each time series. As the Mackey-Glass $τ$ parameter characterizes the time-delay of the system, its increase roughly corresponds to prediction difficulty, shown by the increasing sMAPE trend through the plot. Notably, the LSTM maintains an error that is relatively lower than that of the ESN for all $τ>18$ . However, the LSTM uses explicit memory via input buffering, so it is conjectured that the historical data allows for greater robustness to the varying time series characteristics. The ESN uses only one previous timestep, so its memory is only implicitly retained within its hidden layer. While the ESN tunes well to the $τ$ =17 case and demonstrates greatly reduced effective operations compared to the LSTM, the same set of hyperparameters does not generalize as well to other time series. Further research is motivated in explicit memory buffers versus implicit memory within the network state for trade-offs in single-series forecasting performance, complexity, and generalization capability.
#### Discussion and Opportunities for Further Research
Baseline results for the four v1.0 algorithm track tasks compare the correctness and complexities of various solution types. Compared to ANNs, SNNs and ESNs demonstrate complexity advantages such as smaller footprints, high sparsity, and accumulate rather than multiply-and-accumulate operations. Especially on the motor prediction and chaotic function prediction regression tasks, the SNN and ESN baselines already achieve competitive correctness at lower complexity than the ANN and LSTM counterparts. Further research opportunities in model architectures, data pre-processing and buffering, and training paradigms to achieve greater performance is enabled by the standard framework and tooling provided by NeuroBench.
## System Track Benchmark Framework
While the algorithm track aims to benchmark solutions in a system-independent manner via complexity analysis, the NeuroBench system track aims to evaluate deployed execution time, throughput, and efficiency of systems comprised of an algorithm deployed and tailored to a hardware platform. Previous benchmark studies have examined neuromorphic systems under various applications, including keyword spotting [39, 54], audio and video processing [55], and combinatorial optimization [56, 57]. While these studies have demonstrated neuromorphic system advantages, the benchmark tasks have been unaligned. In order for the hallmarks of neuromorphic hardware to be aptly judged against conventional systems and foster the expansion of neuromorphic solutions, transparent and objective comparisons must be made on standard tasks between sufficiently mature neuromorphic systems head-to-head, as well as against conventional systems.
<details>
<summary>extracted/6132287/figures/system_types.png Details</summary>

### Visual Description
## Diagram: Hardware Scalability Hierarchy (Edge to Multi-Board System)
### Overview
The image illustrates a hierarchical hardware architecture for neuromorphic computing, demonstrating how components scale from a single, standalone edge device to a high-density, multi-board computing system. The diagram uses a color-coded legend to identify specific hardware components across three levels of integration.
### Components/Axes
The diagram utilizes a legend at the bottom to define the hardware components:
* **Black Square:** Sensor
* **Yellow Square:** Flash
* **Red Square:** Supporting microcontroller
* **Light Green Square:** Neuromorphic chip
* **Grey Rectangle:** High-speed link
### Detailed Analysis
#### 1. Edge Device (Left)
* **Description:** A single, small, square-shaped board.
* **Components:**
* 1x Sensor (Black)
* 1x Flash (Yellow)
* 1x Supporting microcontroller (Red)
* 1x Neuromorphic chip (Light Green)
* **Layout:** The components are arranged in a compact, non-grid configuration.
#### 2. Accelerator Board (Center)
* **Description:** A large, rectangular board designed for high-density processing.
* **Components:**
* **Neuromorphic Chips:** A grid of 40 chips (5 rows by 8 columns) in Light Green.
* **Control/Storage:** 1x Supporting microcontroller (Red) and 1x Flash (Yellow) located in the top-left corner of the board.
* **Connectivity:** 5x High-speed links (Grey) positioned along the left vertical edge of the board.
#### 3. Multi-Board System (Right)
* **Description:** A high-performance computing configuration consisting of two vertical stacks.
* **Structure:**
* Two stacks are visible, positioned one above the other.
* Each stack contains 4 Accelerator Boards, totaling 8 Accelerator Boards in the system.
* Each board within the stack maintains the same internal component layout as the standalone Accelerator Board described above.
### Key Observations
* **Scaling Strategy:** The architecture demonstrates a clear modular scaling path: from a single chip (Edge) to a 40-chip board (Accelerator), to a 320-chip system (Multi-Board System, assuming 8 boards x 40 chips).
* **Connectivity Requirement:** The "High-speed link" is absent on the Edge Device but present on the Accelerator Board, indicating that as the system scales, inter-board communication becomes a critical design requirement.
* **Standardization:** The "Supporting microcontroller" and "Flash" components appear to be standard control units that accompany the neuromorphic processing array at both the single-board and multi-board levels.
### Interpretation
This diagram represents a roadmap for neuromorphic hardware deployment.
* **Edge Device:** Likely intended for localized, low-power, real-time inference tasks where data is captured directly by the sensor and processed on-site.
* **Accelerator Board:** Represents a transition to high-performance computing (HPC). The grid-like arrangement of 40 neuromorphic chips suggests a massively parallel processing architecture designed for complex neural network workloads.
* **Multi-Board System:** This configuration implies a data-center or server-grade environment. By stacking multiple accelerator boards, the system achieves massive scale, likely for training large models or handling high-throughput inference tasks. The inclusion of "High-speed links" suggests that the primary engineering challenge at this scale is managing data bandwidth between the boards to prevent bottlenecks.
</details>
Figure 6: Types of neuromorphic systems at various integration scales.
A key challenge for benchmarking neuromorphic hardware is that systems are implemented and deployed at vastly different scales to serve diverse applications, from cloud services (e.g., multi-chip platforms like Loihi [58] and SpiNNaker [59]) to embedded sensing intelligence (e.g., Speck [60] and SNP [61, 62]). This range is visualized in Figure 6. Existing benchmarks for conventional systems have individual focuses across high-performance [63], datacenter-level computing [19], and embedded processing [64]. Thus, rather than pursuing a one-size-fits-all suite of tasks, the goal of the NeuroBench system track is to develop benchmarks at various scales and use cases, under multiple application areas in which both conventional and neuromorphic platforms may compete. The selected v1.0 NeuroBench system track benchmarks represent key commercial application areas for existing systems, and they differ from the tasks in the algorithm track, which are more research-oriented. As benchmark results continue to identify properties of highly effective algorithms and systems, the two tracks will converge to the same selction of tasks that are seen as the most impactful for future progress in the field.
In this section, we present the system track guidelines outlining metrics and tasks, representing collective design between multiple owners and vendors of neuromorphic hardware. Baseline benchmark results for neuromorphic and conventional systems are reported, and further official results will be collected and announced at a regular cadence, akin to the MLPerf suite [65]. As with all other facets of the NeuroBench framework, the system track guidelines will continue to be adapted and extended iteratively as benchmark results are produced and shared. Up-to-date information on the latest benchmarks and official results can be found on the NeuroBench website (https://neurobench.ai/).
### System Track Metrics
In order to be representative of the properties of a deployed system, the system benchmarks, like the algorithm benchmarks, are assessed at the task level for the overall system, as opposed to operation- or kernel-level assessment of individual components. Task-level benchmarks enable straightforward comparison between systems of any type with regard to their abilities to solve problems, and the overall system-level measurement describes the realistic capability and efficiency of a whole solution.
Each individual system benchmark uses task-specific metrics aligned with correctness, timing, and efficiency to measure the system under test (SUT). The following general considerations are applied to each category:
- Correctness – In other system benchmarks, such as the closed category of the MLPerf Inference framework [19], the same trained model is used to benchmark all SUTs, and a correctness threshold is imposed to ensure optimizations such as lower precision do not disrupt task performance.
Due to the tight coupling between an algorithm and its system implementation in many existing neuromorphic hardware solutions, the particular model used to solve a NeuroBench system track benchmark task is unconstrained. Therefore, correctness must be measured to verify the validity of the solution. No correctness thresholds are imposed on submissions, but the benchmark leaderboard will impose tiers of solution correctness on submissions to evaluate accuracy-efficiency trade-offs of system approaches.
- Timing – Depending on the task, timing performance can include measurements of sample throughput or execution time. Individually, the former entails an offline, batched inference benchmark, while the latter aligns with a streaming benchmark, in which one inference does not start until the previous one ends. Together, both throughput and execution time should be reported for tasks in which the SUT runs multiple inferences at any given time, each representing a request which must be responded to within a constrained window. The MLPerf Inference framework has defined widely-adopted general task scenarios corresponding to each of these categories (offline, single-stream, and server, respectively), and the NeuroBench system track will use these scenario guidelines where applicable to maintain consistency and build on conventional frameworks.
In addition, neuromorphic systems are also applied to tasks in which there is no notion of discrete sample throughput or execution, such as for heuristic approximations of intractable problems or operation over a continuous stream of data (e.g., from an event camera). Timing performance should be defined on a per-benchmark basis for such tasks, such as a time-to-solution latency or percentage of execution which exceeds a real-time threshold.
- Efficiency – Conventional system benchmarks such as TOP500 [63] for HPC and MLPerf Inference [19] for deep learning do not require power measurement submission in the main benchmark, instead allowing for separate submissions to an adjacent power track (Green500 [66] and MLPerf Power [67], respectively). Not only has efficiency been usually considered as a second-order metric for conventional systems, it is also notoriously difficult to precisely measure. However, as energy efficiency is a key hallmark of biology and thus is a focus of neuromorphic research, power and energy consumption must be first-order metrics in the NeuroBench system track. Similarly to timing metrics, efficiency metrics should be tailored on a per-benchmark basis, i.e., a real-time always-on processing task may focus on average power, while offline batched systems focusing on high-throughput inference may focus on both peak power and energy per inference.
As neuromorphic systems currently explore a broad range of varied implementation approaches, board-level integration, and developmental stages of hardware, platform diversity poses difficult challenges for completely consistent hardware measurement methodologies (e.g., consistent power meters, chip interfaces, and data loading). Thus, to enable an initial step towards consistency in the system track while ensuring openness, we focus on the development of guidelines for transparent documentation, as they provide the foundation for shared methodology among highly diverse solutions. While there may be differences in how metrics are measured, salient details will be available to contextualize the results, allowing for holistic analysis, and leading the way for future consistency by enforcing transparency.
Benchmark submissions may perform separate runs to report performance and power in order to demonstrate system flexibility (e.g., a ‘performance-mode’ run optimal for execution time and an ‘efficiency-mode’ run optimal for energy), however in all runs, both metrics must be reported.
Importantly for the NeuroBench system track, in measuring timing and efficiency, data pre- and post-processing must be taken into account. Neuromorphic methods will often consume and produce non-standard (e.g., event-based) data modalities, the processing of which may consume a significant amount of the overall execution time and may not be computed on the neuromorphic hardware itself. As many instances of neuromorphic hardware cannot be deployed without such associated processing, it is essential that measurements capture the cost of data processing, which stands in contrast with conventional system benchmarks whose measurements start from pre-processed data [65].
### System Track Benchmarks
Two benchmark specifications for the v1.0 system track are defined in this article, covering embedded to datacenter scales. Full benchmark details are available in the Methods section.
- Acoustic Scene Classification – The acoustic scene classification benchmark challenges systems to classify audio into predefined categories based on the environmental audio context. Such capabilities are key for hearable devices, which can utilize them to automatically adjust sound equalisation profiles, appropriately target microphone denoising, and support active noise cancellation. The application further challenges systems to fulfill technical requirements, such as always-on and real-time operation, and time series processing. Acoustic scenes provide a rich repertoire of features that are necessary for prediction, thus this task is a complement to keyword classification, which mainly focuses on shorter-term features (e.g., phonemes) with a relatively smaller feature repertoire.
The benchmark evaluates the classification capabilities of both neuromorphic systems and conventional computing platforms using datasets from the DCASE challenge [68]. These datasets consist of a myriad of audio recordings from diverse environments, including airports, public parks, and buses, thus providing a comprehensive foundation for testing both application- and system-level performance. The NeuroBench subset of the DCASE dataset includes 41360/16240 train/test samples across four classes (airport, street traffic, bus, park).
The task will be presented under the single-stream task scenario, providing one 1-second sample to the SUT at a time. Classification probability will be sampled to determine the correctness of the prediction. As the NeuroBench system track allows for unconstrained algorithmic implementation, pre-processing and inference metrics should be separately measured and reported together, which differs from prior system benchmarking that only measure inference [39, 64]. Timing results report on-device average execution time per sample. Since the platform diversity of edge-targeted systems poses inherent inconsistencies in efficiency measurement, power should be reported under idle and active contexts, following prior benchmark study methodology [39]. Idle power measures the system prepared for inference with the model loaded, and active power measures the system running pre-processing or inference. The difference between active and idle measurements offers dynamic power, which is used along with execution time to calculate dynamic energy-per-sample.
- QUBO – As a non-ML task, NeuroBench incorporates quadratic unconstrained binary optimization (QUBO). QUBO is a particularly beneficial first optimization task for NeuroBench for two reasons. First, the binary variables are a natural fit for neuromorphic systems with purely binary spike communication. Second, real-world QUBO applications typically feature sparse cost matrices [69] which benefit from the sparse synaptic connectivity and execution that neuromorphic systems are often optimized for [70]. The initial set of QUBO workloads in NeuroBench searches for the maximum independent (i.e. unconnected) set of nodes in graphs, a task that has wide applications across industry and academia, such as resource allocation in wireless networks, portfolio optimization, and task scheduling [71].
NeuroBench provides a QUBO generator that can uniquely specify each workload by three specific parameters provided by the benchmark: the number of graph nodes, the density of graph connections, and a random seed. The generator provides a large dataset for reliable statistics and allows scaling from modest workloads for small-scale and prototype systems to large workloads for larger-scale systems. The graph sizes specified by the benchmark increase in a pseudo-geometric progression (10, 25, 50, 100, 250, …), and submissions are encouraged to extend the problem size to the limits of the SUT. Graph density ranges from 1% to 30%, in order to show the relationship between the number of connections and SUT power. 5 random seeds for each setting should be tested.
Optimization algorithms use heuristic methods to iteratively refine approximate solutions to intractable problems that cannot be completely solved. The QUBO benchmark thus measures solution optimality and energy consumption after fixed, pre-set runtimes, removing any timing measurement. Solution optimality is defined as BKS-Gap, a relative gap between the current SUT’s solution compared against the best-known solution (BKS) to the same problem found using a high-powered solver with a long runtime.
### Baseline Results
Baseline results for each of the two system track benchmarks are provided for a mature neuromorphic system against a conventional platform. Like for the algorithm track baseline results, the system track baselines are intended to snapshot the solution space and provide starting points for the task leaderboards. Further details on each baseline system are available in the Methods section, and in-depth system documentation for the Xylo ASC baseline and CPU/Loihi 2 QUBO baselines are provided by Ke et al. [72] and Pierro et al. [57], respectively.
#### Acoustic Scene Classification
For the acoustic scene classification task, two baseline embedded systems are reported:
- CPU – The CPU baseline is an Arduino Nano 33 BLE, which uses an ARM Cortex M4 microcontroller for both pre-processing and inference. The digital audio sample is pre-processed using Mel-filterbank energies (MFE), and inference uses a conventional CNN. Execution time is measured using on-chip timers, and power is measured using total system power.
- Xylo – The Synsense Xylo [73] neuromorphic baseline uses a feed-forward SNN with multiple synaptic time constants [54]. As the system is intended for continuous, real-time audio processing, the board uses an analog front end to pre-process analog audio signals directly from a microphone into spikes for the digital inference engine. To conform with a digital benchmark dataset, a simulator of the analog pre-processor generates spikes, which are routed to the inference module. Execution time and power are measured using on-board instruments.
Baseline Accuracy Execution Time (ms) Idle Power (mW) Active Power (mW) Dynamic Power (mW) Dynamic Energy (mJ/inf) CPU pre-process 79.64% 43 79.40 100.72 21.32 0.917 (system-wise) inference 45 79.40 100.15 20.75 0.934 Xylo pre-process* 79.90% - 0.00017* 0.015* 0.015* 0.015* (component-wise) inference 84 0.351 0.692 0.341 0.028
Table 6: Baseline results for the acoustic scene classification task. Pre-processing of the Xylo (marked with an asterisk *) measures power of the analog pre-processor in real-time relative to the audio data, whereas other measurements are of digital components processing digital data on-hand. Idle Xylo pre-processing measures silence and active measures test data audio played to the device, and energy is measured over the sample duration of 1 second. CPU power is measured over the full Arduino system, as it does not have on-board power instrumentation, while Xylo power measures power consumed by the Xylo Audio 2 ASIC only. Dynamic power and energy provide proper comparison between the systems, and idle and active measurements are provided for transparency.
Table 6 lists baseline results for the neuromorphic Xylo system against an Arduino system. Compared to prior neuromorphic audio system benchmarking [39], which takes a server-class CPU as a point of comparison, we adopt a fairer approach by focusing on low-power edge application and comparing against an Arduino embedded microprocessor. At comparable inference accuracy, Xylo exhibits $60.9×$ less dynamic inference power and $33.4×$ less dynamic inference energy consumption than the Arduino.
#### QUBO
<details>
<summary>x4.png Details</summary>

### Visual Description
## Multi-Panel Line Chart: Performance Gap vs. Number of Variables
### Overview
This image displays a grid of six line charts, arranged in a 2x3 layout. Each chart plots the "% Gap from Best Known Solution" (Y-axis) against the "Number of Variables" (X-axis) for a specific "Timeout" duration. There are three distinct data series represented by colored lines (Blue, Orange, and Green), each accompanied by a shaded region representing uncertainty or variance (confidence interval).
There is no explicit legend provided in the image to identify the specific algorithms or methods corresponding to the Blue, Orange, and Green lines.
### Components/Axes
* **Layout:** 6 subplots arranged in 2 columns and 3 rows.
* **X-Axis:** "Number of Variables" (Logarithmic scale). Ticks are present at $10^1$ and $10^3$, with intermediate grid lines representing logarithmic intervals.
* **Y-Axis:** "% Gap from Best Known Solution".
* **Top Row (Timeouts $10^{-3}$ s, $10^{-2}$ s):** Scale ranges from 0 to 100.
* **Middle Row (Timeouts $10^{-1}$ s, $1$ s):** Scale ranges from 0 to 20.
* **Bottom Row (Timeouts $10^1$ s, $10^2$ s):** Scale ranges from 0 to 10.
* **Titles:** Each subplot is titled with its respective timeout duration:
* Top-Left: `Timeout @ 10⁻³ s`
* Top-Right: `Timeout @ 10⁻² s`
* Middle-Left: `Timeout @ 10⁻¹ s`
* Middle-Right: `Timeout @ 1 s`
* Bottom-Left: `Timeout @ 10¹ s`
* Bottom-Right: `Timeout @ 10² s`
---
### Detailed Analysis
#### Row 1: High-Speed Timeouts ($10^{-3}$ s and $10^{-2}$ s)
* **Trend:** All series show a sharp upward trend as the number of variables increases.
* **Top-Left ($10^{-3}$ s):**
* **Orange:** Rises steeply from 0 at $10^1$ to 100 at $10^2$, remaining at 100 thereafter.
* **Green:** Remains at 0 until $10^2$, then rises sharply to 100 at $5 \times 10^2$.
* **Blue:** Remains at 0 until $10^2$, then rises to approximately 25 at $5 \times 10^2$, where it plateaus.
* **Top-Right ($10^{-2}$ s):**
* **Orange:** Rises to ~25 at $5 \times 10^2$, then jumps to 100 at $10^3$.
* **Green:** Rises to ~15 at $5 \times 10^2$, then jumps to 100 at $10^3$.
* **Blue:** Rises to ~20 at $5 \times 10^2$, remaining flat at $10^3$.
#### Row 2: Moderate Timeouts ($10^{-1}$ s and $1$ s)
* **Trend:** The gap is significantly lower than the top row. The curves show a "hump" shape or a plateau after $5 \times 10^2$ variables.
* **Middle-Left ($10^{-1}$ s):**
* **Orange:** Shows the highest gap, reaching ~18 at $10^3$.
* **Blue:** Reaches ~14 at $5 \times 10^2$, then dips slightly to ~12 at $10^3$.
* **Green:** Reaches ~10 at $5 \times 10^2$ and stays flat.
* **Middle-Right ($1$ s):**
* **Blue:** Reaches ~10 at $5 \times 10^2$, then dips to ~8 at $10^3$.
* **Orange:** Reaches ~10 at $10^3$.
* **Green:** Reaches ~7 at $5 \times 10^2$, then dips slightly.
#### Row 3: Long Timeouts ($10^1$ s and $10^2$ s)
* **Trend:** The gap is minimized. The curves show a clear separation between the three series.
* **Bottom-Left ($10^1$ s):**
* **Blue:** Highest gap, peaking at ~8.5 at $5 \times 10^2$, then dropping to ~7.5 at $10^3$.
* **Orange:** Reaches ~5.5 at $10^3$.
* **Green:** Lowest gap, reaching ~4 at $10^3$.
* **Bottom-Right ($10^2$ s):**
* **Blue:** Highest gap, reaching ~6 at $10^3$.
* **Orange:** Reaches ~2.5 at $10^3$.
* **Green:** Lowest gap, reaching ~2 at $10^3$.
---
### Key Observations
1. **Inverse Relationship:** As the "Timeout" duration increases (moving from top-left to bottom-right), the "% Gap from Best Known Solution" decreases significantly.
2. **Scaling Difficulty:** As the "Number of Variables" increases, the gap generally increases, indicating that the problem becomes harder to solve optimally as complexity grows.
3. **Algorithm Performance:**
* **Green Series:** Generally performs best (lowest gap) in the longer timeout scenarios (bottom row).
* **Blue Series:** Performs best in the shortest timeout scenarios (top row) but performs worst (highest gap) in the longest timeout scenarios (bottom row).
* **Orange Series:** Often occupies the middle ground or performs worst in the shortest timeouts.
### Interpretation
The data demonstrates a classic trade-off in computational optimization algorithms. The three series likely represent different algorithms or heuristics.
* **The "Blue" algorithm** appears to be a "fast-start" or "greedy" approach. It performs well when time is extremely limited ($10^{-3}$ s), likely finding a decent solution quickly. However, it fails to improve significantly with more time, resulting in a higher gap compared to the others when given $10^1$ or $10^2$ seconds.
* **The "Green" algorithm** appears to be a more robust or "thorough" approach. It struggles significantly when time is extremely limited (high gap in top row) but scales much better as time is provided, eventually achieving the lowest gap from the best-known solution in the longest timeout scenarios.
* **The "Orange" algorithm** represents a middle-ground approach, though it is notably outperformed by the Green algorithm in the long-timeout scenarios.
The "hump" or dip observed in the middle and bottom rows (where the gap decreases slightly at $10^3$ variables compared to $5 \times 10^2$) suggests that for certain algorithms, the problem structure at $10^3$ variables might be easier to approximate than at $5 \times 10^2$, or the algorithm's heuristic is better suited to that specific scale.
</details>
<details>
<summary>x5.png Details</summary>

### Visual Description
## Legend: Data Series Identification
### Overview
The image is a standalone legend box, likely extracted from a larger technical chart or graph. It serves to define three distinct data series using color-coded horizontal line segments paired with text labels. The legend is enclosed within a single, rounded-corner rectangular border.
### Components/Axes
The legend contains three distinct entries arranged horizontally from left to right:
1. **Left Entry**: An orange horizontal line segment followed by the text "SA".
2. **Center Entry**: A teal/green horizontal line segment followed by the text "TABU".
3. **Right Entry**: A blue horizontal line segment followed by the text "Loihi".
### Detailed Analysis
* **Spatial Grounding**: The legend is centered horizontally within the image frame. The three entries are spaced at roughly equal intervals.
* **Color/Label Mapping**:
* **SA**: Represented by an orange line.
* **TABU**: Represented by a teal/green line.
* **Loihi**: Represented by a blue line.
* **Typography**: The text uses a serif typeface. The lines are of uniform length and thickness across all three entries.
### Key Observations
* The legend uses a standard convention of color-coding to distinguish between three different methodologies, algorithms, or hardware platforms.
* The visual consistency of the line segments suggests they represent the same data type (e.g., lines on a line graph) across the three categories.
### Interpretation
This legend provides critical context for a comparative study, likely found in a research paper or technical report. Based on the labels, the data likely represents a performance comparison:
* **SA**: Almost certainly refers to **Simulated Annealing**, a probabilistic technique for approximating the global optimum of a given function.
* **TABU**: Almost certainly refers to **Tabu Search**, a metaheuristic search method used for mathematical optimization.
* **Loihi**: This is a specific reference to **Intel's Loihi**, a neuromorphic research chip designed to mimic biological neural networks.
**Synthesis**: The chart this legend belongs to is likely evaluating the performance (such as execution time, energy efficiency, or solution quality) of two classical optimization algorithms (SA and TABU) against a neuromorphic hardware implementation (Loihi). The presence of "Loihi" strongly suggests the document is investigating the efficacy of neuromorphic computing in solving combinatorial optimization problems compared to traditional CPU-based algorithms.
</details>
Figure 7: Percentage gap from the best known solution (BKS-Gap%) for the QUBO workloads with QUBO matrices at $15\$ density (lower is better). Results are shown for different timeouts of the QUBO solvers. Figure taken, with permission, from Pierro et al. [57].
<details>
<summary>x6.png Details</summary>

### Visual Description
## Line Charts: Average Power Rate vs. Number of Variables by Density
### Overview
The image displays three side-by-side line charts comparing the "Average Power Rate [W]" against the "Number of Variables" across three different "Density" configurations: 5%, 15%, and 30%. The charts utilize a logarithmic scale for both the X and Y axes. The data indicates that power consumption is stratified into distinct tiers, with some tiers remaining constant regardless of the number of variables, while others show a slight upward trend as the number of variables increases.
### Components/Axes
* **Layout:** Three identical chart frames arranged horizontally.
* **X-Axis:** Labeled "Number of Variables" at the bottom of each chart. The scale is logarithmic, ranging from $10^1$ (10) to $10^3$ (1000).
* **Y-Axis:** Labeled "Average Power Rate [W]" on the far left. The scale is logarithmic, ranging from $10^{-1}$ (0.1) to $10^2$ (100).
* **Titles:**
* Left Chart: "Density 5%"
* Center Chart: "Density 15%"
* Right Chart: "Density 30%"
* **Data Series (No legend provided; identified by visual characteristics):**
* **Tier 1 (High Power):**
* Dark Teal (Solid line, circle markers): Top-most line.
* Orange (Solid line, circle markers): Second from top.
* Light Orange (Dashed line, circle markers): Third from top.
* **Tier 2 (Medium Power):**
* Green (Dashed line, circle markers): Fourth from top.
* **Tier 3 (Low Power):**
* Dark Blue (Solid line, circle markers): Fifth from top.
* Light Blue (Dashed line, circle markers): Sixth from top.
* Faded Blue (Dashed line, circle markers): Bottom-most line.
### Detailed Analysis
The data patterns are consistent across all three density charts (5%, 15%, and 30%).
**Trend Verification:**
* **Tier 1 & 2 (Top four lines):** These lines (Dark Teal, Orange, Light Orange, Green) exhibit a flat, horizontal trend. The power rate remains constant regardless of the increase in the number of variables.
* **Tier 3 (Bottom three lines):** These lines (Dark Blue, Light Blue, Faded Blue) exhibit a positive slope, particularly visible between $10^2$ and $10^3$ variables, indicating that power consumption increases as the number of variables increases.
**Approximate Data Points (Consistent across all three charts):**
* **Dark Teal (Solid):** Remains constant at approximately $10^2$ W (100 W).
* **Orange (Solid):** Remains constant at approximately $70-80$ W.
* **Light Orange (Dashed):** Remains constant at approximately $60$ W.
* **Green (Dashed):** Remains constant at approximately $13$ W.
* **Dark Blue (Solid):** Starts at $\approx 2$ W at $10^1$ variables, rising to $\approx 3$ W at $10^3$ variables.
* **Light Blue (Dashed):** Starts at $\approx 1.3$ W at $10^1$ variables, rising to $\approx 2$ W at $10^3$ variables.
* **Faded Blue (Dashed):** Starts at $\approx 0.8$ W at $10^1$ variables, rising to $\approx 1.2$ W at $10^3$ variables.
### Key Observations
* **Density Invariance:** The "Density" parameter (5%, 15%, 30%) appears to have no significant impact on the power consumption profiles, as the charts are visually identical.
* **Power Stratification:** The system power consumption is clearly divided into three distinct operational tiers. The high-power tiers (above 10W) appear to represent fixed overhead or background processes that do not scale with the workload (number of variables).
* **Scaling Behavior:** Only the low-power tier (below 3W) shows sensitivity to the "Number of Variables." This suggests that the components represented by the blue lines are the only ones performing work that scales with the input size.
### Interpretation
The data suggests a system architecture where the majority of power consumption is driven by static, non-scaling components (the high-power tiers).
* **Fixed Overhead:** The fact that the top four lines are flat suggests these represent system-level components (e.g., cooling, idle CPU cores, memory controllers, or background services) that consume a constant amount of power regardless of the computational load.
* **Workload Scaling:** The bottom three lines represent the only components that scale with the "Number of Variables." This is likely the primary compute unit or the specific algorithm being tested.
* **Efficiency Implications:** If the goal is to reduce power consumption, optimizing the "Number of Variables" will yield diminishing returns because the high-power, non-scaling components dominate the total power budget. To significantly reduce total power, one would need to address the high-power tiers (the 10W-100W range) rather than the scaling components.
* **Anomaly/Outlier:** There is a slight dip in the Faded Blue line at $10^2$ variables in the "Density 30%" chart, which is not present in the other charts. This could be a minor measurement artifact or a specific system state transition at that density.
</details>
<details>
<summary>x7.png Details</summary>

### Visual Description
## Legend: Multi-Series Data Identification
### Overview
The image is a standalone legend component, likely extracted from a larger technical chart or graph. It provides a key to identify three distinct data series using color-coded line segments.
### Components
The legend is enclosed within a rounded rectangular border. It contains three entries arranged horizontally from left to right:
1. **Gold/Mustard Line Segment**: Labeled "SA"
2. **Teal/Green Line Segment**: Labeled "TABU"
3. **Blue Line Segment**: Labeled "Loihi"
### Content Details
* **SA**: Represented by a gold/mustard-colored horizontal line.
* **TABU**: Represented by a teal/green-colored horizontal line.
* **Loihi**: Represented by a blue-colored horizontal line.
### Key Observations
* The legend uses distinct, high-contrast colors to differentiate between the three series.
* The layout is linear and horizontal, suggesting it was likely positioned at the top or bottom of a chart.
### Interpretation
The labels provided in this legend strongly suggest a comparison of optimization algorithms or computing architectures:
* **SA**: Likely refers to **Simulated Annealing**, a probabilistic technique for approximating the global optimum of a given function.
* **TABU**: Likely refers to **Tabu Search**, a metaheuristic search method used for mathematical optimization.
* **Loihi**: Refers to **Intel's Loihi**, a neuromorphic research chip designed to mimic biological brain structures for efficient computing.
The data associated with this legend likely compares the performance, energy efficiency, or convergence speed of these three distinct approaches (two classical optimization algorithms vs. one neuromorphic hardware platform).
</details>
<details>
<summary>x8.png Details</summary>

### Visual Description
## Legend: Resource Utilization Chart
### Overview
The image is a standalone legend box, likely extracted from a technical line graph. It provides a key for three distinct data series, utilizing line stroke patterns to differentiate between categories.
### Components
The legend is enclosed within a light gray, rounded rectangular border. The elements are arranged horizontally from left to right:
* **Left:** A dashed line symbol followed by the text "Compute".
* **Center:** A dotted line symbol followed by the text "Memory".
* **Right:** A solid line symbol followed by the text "Total".
### Detailed Analysis
The legend defines the visual encoding for the data series in the associated chart:
* **Compute:** Represented by a dashed line (long, distinct segments).
* **Memory:** Represented by a dotted line (short, frequent dots).
* **Total:** Represented by a solid, continuous line.
### Key Observations
The legend relies exclusively on line stroke patterns (dashed, dotted, solid) to differentiate data series rather than color. This is a robust design choice, ensuring the chart remains readable in grayscale, monochrome printing, or for users with color vision deficiencies.
### Interpretation
This legend is characteristic of system performance monitoring or resource allocation analysis. It implies that the underlying chart tracks three metrics: "Compute" (likely CPU or processing power), "Memory" (RAM or storage usage), and "Total" (the aggregate of the two or a combined system load metric).
Given the labels, it is highly probable that the "Total" line represents the sum of the "Compute" and "Memory" values, or a broader system metric that encompasses both. The use of distinct line styles suggests the chart plots these values against a continuous variable, such as time or workload intensity.
</details>
Figure 8: Power consumption of the QUBO solvers running simulated annealing (SA) or TABU search on CPU, and the parallelized version of simulated annealing on Loihi 2. The CPU solvers require up to $37×$ more power than the neuromorphic algorithm on Loihi 2. Since the QUBO workloads were run for a fixed timeout, differences in power consumption are equivalent to differences in energy consumption of the processors. Figure taken, with permission, from Pierro et al. [57].
Three baselines are measured for the QUBO benchmark:
- Simulated Annealing (SA) – The simulated annealing solver uses Markov chain Monte Carlo (MCMC) sampling to probabilistically explore the search space.
- Tabu Search (TABU) – Tabu search solvers maintain and iterate a list of prohibited actions in order to prevent the search from remaining in local minima or revisiting states. Both the TABU and SA solver baselines use the D-Wave Samplers library [74] on an Intel Core i9-7920X desktop-class CPU, with power measured using Intel SoC Watch.
- Loihi 2 – The Loihi 2 neuromorphic system solver uses an SNN formulation of the simulated annealing algorithm which enables solving via neural dynamics, and parallelization via stochastic refractory periods. The baseline is implemented on one Loihi 2 chip on the 8-chip Kapoho Point board, and internal power instrumentation measures all compute and memory components of the chip.
Figure 7 shows the optimality reached by the CPU- and Loihi 2-based solvers after different timeouts. For tight time constraints, at $10^-2$ seconds timeout or less, Loihi 2 finds feasible solutions to workloads $4×$ larger than the CPU. But for timeout lasting $10$ seconds or longer, the CPU running TABU provides the lowest BKS-Gap, incentivizing algorithmic advances for neuromorphic optimization systems. Figure 8 illustrates the power consumed during runtime. Across the workloads, the Loihi 2 solver requires $37.24×$ less power compared to the best CPU solver.
#### Discussion and Future Work
The initial baselines for the v1.0 system track compare correctness, timing, and efficiency of neuromorphic systems against conventional CPU systems in domains of both audio classification and optimization. Against mature, commercially-developed CPU systems, for both edge and server use cases, the neuromorphic systems show strong advantages in general efficiency, as well as further promises in terms of timing and correctness.
In future NeuroBench iterations, the system track benchmarks can be unified under common tooling, similar to the algorithm track. Software toolchains such as Lava [36], Fugu [37], SPyNNaker [75], and Samna [76], among others, have been developed to interface with specific hardware platforms. Many of the stacks are built with general paradigms to support extension to any backend, and the community is actively moving towards developing standards for deployment tools. The current v1.0 benchmark specifications allow for open algorithm and software design in order to demonstrate fully optimized performance for neuromorphic systems. As standards mature in the future, a core focus of the NeuroBench system track is to introduce a closed-algorithm benchmarking category that leverages the recently proposed NIR model description framework [77] as a general, cross-platform tool for benchmarking key workloads of interest across many different platforms.
## Discussion
Benchmarking neuromorphic computing has faced challenges stemming from the diversity of neuromorphic approaches, the range of implementation and deployment tools, and rapid research evolution. NeuroBench addresses these challenges as a framework for the inclusive, actionable, and iterative benchmarking of neuromorphic solutions, by including novel tasks and metrics, open-source and extendable harness tooling, and facilitating systematic growth via community collaboration. NeuroBench is supported and developed by a broad community of neuromorphic researchers to be a standard, agreed-upon benchmarking framework for neuromorphic technology.
Initial NeuroBench benchmarks span applications across domains of continual learning, computer vision, sensorimotor prediction, and time-series forecasting, as well as system implementations for audio and optimization settings. Baselines for each benchmark of the complete v1.0 algorithm track demonstrate the utility and validity of the metric framework, and offer a starting point to further algorithmic research in model architecture and training for greater performance and lowered complexity. The system track v1.0 benchmark baselines similarly provide the groundwork for both conventional and head-to-head comparisons of deployed hardware systems. Via collaboratively-defined measurement methodology and task specifications, the shared system track guidelines will be modified and extended to define standard protocols for continuous-time execution, analog circuit implementations, and other exploratory platforms in simulation stages, such as memristive hardware.
Another important direction for NeuroBench is towards closed-loop benchmarks [15, 78]. Biological systems excel in interacting with dynamic environments, demonstrating high energy efficiency, real-time reaction, and versatility. As such, embodied intelligence with adaptive sensory and action capabilities are of interest to neuromorphic research. In closed-loop scenarios, the objective is to sense and act within an environment to complete a task, rather than to statically process a frozen dataset, thus the benchmark harness infrastructure and measurement protocols will be extended to facilitate such benchmarks.
All future NeuroBench expansion will be informed by collected results and continue to be driven by the interests and development of the broader community.
## Methods
This section outlines details and specifications of the benchmark metrics, tasks, and baselines.
### Algorithm Track Metrics
NeuroBench includes correctness and complexity metrics, the latter of which is divided in static and workload metrics. Static metrics do not depend on the model inference and input data, while the workload metrics do. Note that the defined metrics reflect only the model and model execution. Data pre-processors and post-processors are not taken into account in the v1.0 algorithm track results.
#### Footprint
The footprint metric reflects the memory footprint a model. It is distinct from execution memory, which may incur further usage, e.g. to store activations. It is computed for a model by accumulating the sizes of the model’s parameters and buffers, in bytes. Parameters store the model synaptic weights, and buffers include other inference memory requirements, such as the internal states of recurrent or spiking layers and buffers of recent input data, if the model must record data for input binning. Considering $n$ parameters, each requiring $p_i$ bytes, and $b$ buffers of size $q_j$ , the total model footprint is $∑_i=0^np_i+∑_j=0^bq_j$ .
#### Model Execution Rate
Execution rate is a numeric which is not directly computed by the harness, but should be reported by the user. The numeric reflects the real-time correlation of the rate at which the model computes input data. If the model processes input with a temporal stride of $t$ seconds, then the rate should be reported as $t^-1$ Hz. Note the distinction between stride and bin window - input can be binned in overlapping windows, but execution rate depends on the temporal stride of window processing. As an example, a model may use 50 ms windows of input and compute every 10 ms, which would give an execution rate of 100 Hz.
This numeric is currently not well-defined for models operating under event-based or continuous-time contexts. These limitations will be addressed in future benchmark versions.
#### Connection sparsity
The parameter matrices of each layer $l$ in a model, representing synaptic weights, are collected, and the number of zero weights $m_l$ and total weights $n_l$ are aggregated, with the connection sparsity defined as $\frac{∑_lm_l}{∑_ln_l}$ .
#### Activation sparsity
Activation sparsity is computed after the inference phase. The sparsity is calculated by accumulating the number of zero activations ( $z$ ), over all neuron layers ( $l$ ), timesteps ( $t$ ), and input samples ( $i$ ) and dividing by the total number of neurons ( $N$ ), $\frac{∑_l∑_t∑_iz_l,t^i}{∑_t∑_l∑_iN_l,t^i}$ . The outputs of ReLU functions and spikes from spiking neurons are considered activations.
#### Synaptic operations
Synaptic operations are the multiplication of weights by activation or input data, and are calculated using the inputs and weights of connection layers (e.g., torch.nn.Linear and torch.nn.Conv2d). Effective synaptic operations are operations where a non-zero weight is multiplied by a non-zero activation. Effective operations are further divided into multiply-accumulates (MACs), and accumulates (ACs), where accumulates correlate with activations or input data only containing values of [-1, 0, 1], and multiply-accumulates cover all other cases. The reported number of synaptic operations is the average number of synaptic operations required per model execution, the rate of which is defined by the model execution rate metric.
The number of effective synaptic operations is computed by performing the forward pass of a layer and counting the number of operations in which there is no zero multiplication. Practically, this is implemented in the harness by setting all non-zero weights in the layer and all the non-zero activations to 1, then performing the forward pass and summing the output to give the number of synaptic operations.
The number of dense synaptic operations is computed in a similar fashion, by setting all weights and activations to 1 and accumulating the output of the forward pass. Biases are not taken into account in the calculation of the synaptic operations, as they are added after weight multiplications and accumulation.
Note that processing of activations before the connection layer, for instance using batch normalization, can transform sparse activations into dense input at the connection layer, which will lead to high effective synaptic operations despite high activation sparsity. Furthermore, such processing can transform binary activations to non-binary data, causing effective operations to be MACs rather than ACs. When deployed to neuromorphic hardware, such algorithms that normalize activations before multiplication with synaptic weights may lose the benefits of sparse operation, e.g., an SNN with normalization following each spiking layer would require dense MAC weight calculation, no matter how few spikes were generated.
In some cases, algorithm execution may have distinct temporal sections of higher and lower synaptic operations, such as during initial caching versus continuous inference. For such algorithms, benchmark users may choose to distinguish synaptic operations and other complexity measurements between execution sections.
### Algorithm Track Benchmark Tasks
#### Keyword FSCIL
Few-shot Class-Incremental Learning, FSCIL, is an established benchmark task setting in the computer vision domain [26]. It can be defined as follows: a base session with fixed classes, each with abundant training data, is used to train an initial model. Then, successive incremental training sessions introduce new classes in a few-shot learning scenario. In each session, only the current session classes are available to the model for training. After each incremental training session, the model is evaluated on all previously seen classes, including the base classes. Therefore, the model has to learn new classes while retaining knowledge about the previously learned ones.
Formally, for $M$ -step FSCIL, where $M$ is the total number of incremental sessions, each training session uses a support dataset $D^(t)$ , $t∈[0,M]$ to train new classes on. ${L^(t)}$ is the set of classes of the $t$ -th session where $∀ i,j$ where $i≠ j$ , $L^(i)∩ L^(j)=∅$ , meaning each training session uses a unique set of classes. $D^(0)$ and $L^(0)$ are the base class training data and set of base classes, respectively, $D^(1)$ and $L^(1)$ represent the first incremental session set, and so on. At session $t$ , only $D^(t)$ is available for training, and for $t>0$ , $D^(t)$ contains a fixed number of classes ( $N$ ) with few samples per class ( $K$ ). This form of FSCIL is therefore named $N$ -way $K$ -shot FSCIL. At the end of each session $t$ , model accuracy is reported on the test samples of all previously seen classes $\{L^(0)∪ L^(1)∪...∪ L^(t)\}$ .
For the Keyword FSCIL task, classes in the base set ( $L^(0)$ ) have 700 samples each, with a fixed train/validation/test sample split of 500/100/100. All classes within incremental sessions have 200 samples per word, with a fixed train/test split of 100/100. Of the 100 training samples, 5 are randomly selected for few-shot learning (each session is 10-way, 5-shot). The inclusion of 200 samples allows for increasing learning up to 100 samples.
NeuroBench proposes an audio keyword classification version of the FSCIL task, which to the best of our knowledge is the first of its kind. This novel task is established by selecting a subset of the words and languages from the Multilingual Spoken Word Corpus (MSWC) [21] dataset. The FSCIL task consists of a multilingual set of 100 base classes and 10 incremental sessions of 10 classes each, for a final total of 200 learned classes. Fifteen languages are represented: the base classes are composed of a set of five base languages with 20 words each, and each of the ten incremental sessions contains 10 words from a distinct language. The languages were chosen based on data availability within the MSWC dataset. The top five languages with the greatest number of potential words (words with enough data samples) are used as the base class languages, while the next ten languages with the greatest numbers are the incremental classes. The base languages are English, German, Catalan, French and Kinyarwada. Incremental languages are Persian, Spanish, Russian, Welsh, Italian, Basque, Polish, Esparanto, Portuguese and Dutch. The order of languages presented in the incremental sessions are randomized, but each incremental session will represent exactly one new language.
For each language, the longest length words (that had the appropriate number of samples) were selected to allow for rich and robust temporal features to be learned. Next to the richness of longer words, there are practical considerations for this choice. The MSWC dataset normalizes all samples to a duration of 1 second, centered around the 0.5 seconds mark. For shorter words, this means that the data needs to be zero-padded on both sides to fill the entire duration. Longest-length words are likely to fill the complete sample and reduce zero-padding, which is also useful in scenarios in which algorithms seek to classify words before the sample has completed [79]. Furthermore, common keyword spotting solutions, such as Ok Google, Alexa, and Hey Siri, use multi-syllable wake-phrases to assist in accurate word classification. Using shorter keyword subsets can lead to greater challenge in both base language training and continual class learning. The full list of chosen words for the task presented in this paper (longest words), as well as other potential subsets of short words for the same languages, can be found with dataset documentation through the harness.
Within each language, words showing great similarity in phonics and meaning are not included (e.g. l’amendement and amendements in French). Across different, but related languages, words with similar pronunciation and meaning were not included as well (e.g., university, universität and universitat in English, German and Catalan).
The subset of MSWC used for this FSCIL task is significantly smaller in size (630MB) compared to the full MSWC datatset (124GB), and subset download details can be found in the harness.
#### Event Camera Object Detection
The task of object detection using event camera data involves identifying bounding boxes of objects belonging to multiple predetermined classes in an event stream. The dataset is the Prophesee 1 Megapixel automotive detection dataset [22], which is one of the largest and highest-resolution event-camera detection datasets currently available. The performance of the task is defined by the COCO mean average precision (mAP) metric [28], a metric that is commonly used for the evaluation of object detection algorithms. Only three out of the seven available object classes within the dataset are used due to limited sample availability in the dataset, which matches prior work [22].
COCO mAP is calculated using the intersection over union (IoU, Equation 1) of the bounding boxes produced by the model against ground-truth boxes. Here, $A$ and $B$ refer to bounding boxes, and the intersection and union consider the overlapping area and the area covered by both boxes, respectively. The IoU is compared against 10 thresholds between 0.50 and 0.95, with a step size of 0.05. For each threshold, precision is calculated (Equation 2) with True Positives (TP) and False Positives (FP) determined by whether the IoU meets the threshold or not, respectively. The mAP is calculated as the averaged precision over all thresholds for each class, which is further averaged over all classes to produce the final result.
$$
IoU(A,B)=\frac{|A∩ B|}{|A∪ B|} \tag{1}
$$
$$
Precision(TP,FP)=\frac{∑ TP}{∑ TP+∑ FP} \tag{2}
$$
Note that in the dataset, labels are generated from images from an RGB camera. Due to the nature of event cameras, objects which are still at the start of a recording sequence have no generated events and cannot be detected. Therefore, labels within the first 0.5 seconds of each sequence are not taken into account. Furthermore, as the RGB camera used for labeling has a higher resolution than the event camera, not all objects which appear in the RGB image are recognizable from the generated events. Thus, objects with a diagonal of less than 60 pixels are also not considered. The dataset and metric measurement is implemented using the Prophesee Metavision software [80].
#### Non-human Primate Motor Prediction
The non-human primate motor prediction task involves predictive modeling of two-dimensional fingertip velocity, given neural motor cortex data. The six sessions used for the benchmark comprise three recording sessions each from two non-human primates (NHP Indy and NHP Loco) such that the chosen sessions approximately span the entire duration of the experiment [23] (several months). The specific sessions used are indy_20170131_02, indy_20160630_01, indy_20160622_01, loco_20170301_05, loco_20170215_02, and loco_20170210_03. Each of these sessions consists of one day of experiments, during which multiple reaches are recorded. During each reach, a target position is displayed, which the NHP needs to localize and touch with its finger. Once the NHP touches the correct target for the current reach, the next reach is instantiated, showing a new target position. The data contains sensorimotor cortex recordings from 96 channels for the recordings of the first NHP (Indy), while 192 channels were used for the second NHP (Loco), and was gathered and labeled at a frequency of 250Hz. Two-dimensional position data of the NHP fingertip during its reaches is provided in the dataset, and these are translated into $X$ and $Y$ velocity ground-truth labels using discrete derivatives [81].
Each session is segmented into individual reaches based on the target position for the NHP to touch. The data in each session is split such that the initial $75\$ of reaches are used for training and validation, and the remaining $25\$ of reaches are test data. The user can choose how to utilize the training and validation split for their particular method.
During evaluation, the coefficient of determination ( $R^2$ , Equation 3) for the $X$ and $Y$ velocities are averaged to report the correctness score for each session, where $n$ is the number of labeled points in the test split of the session, $y_i$ is the ground-truth velocity, $\hat{y_i}$ is the predicted velocity, and $\bar{y}$ is the mean of the ground-truth velocities. The $R^2$ from sessions for each NHP are averaged, producing two final correctness scores.
$$
R^2=1-\frac{∑_i=1^n(y_i-\hat{y_i})^2}{∑_i=1^n(y_i-\bar
{y})^2} \tag{3}
$$
#### Chaotic Function Prediction
The chaotic function prediction is another sequence-to-sequence problem. Given an input sequence generated from a one-dimensional Mackey-Glass function, the task is to predict the future values of the same function. The dataset used for this task is synthetically generated, following the Mackey-Glass differential equation [24] (Equation 4), which is integrated and discretized with a timestep of $Δ t$ . The time series generated by this differential equation is a function of the Mackey-Glass parameters $n$ , $β$ , $γ$ and $τ$ . Adhering to standard parameters [29], the values used for $n$ , $β$ , and $γ$ are 10, 0.2 and 0.1 respectively. $τ$ is varied between 17 (a standard value) and 30, leading to 14 time series which vary greatly in dynamics and can be used to analyze the generalization of predictive models.
Each value of $τ$ is associated with a Lyapunov time, the expected predictability timescale for chaos [33], which is used as the time unit for each series. To calculate the overall Lyapunov time for each value of $τ$ , we average the Lyapunov times of 30,000 generated time series of 2,000 timesteps, with $Δ t=1.0$ , each with a randomly chosen initial condition. All time series and Lyapunov times were generated and estimated using the JiTCDDE library [82]. For each final time series used for benchmarking, initial conditions are a point randomly chosen along the series. The Lyapunov time and initial condition $x_0$ for each of the 14 final time series are provided in Table 7.
$$
\frac{dx}{dt}=\frac{β x(t-τ)}{1+x(t-τ)^n}-γ x(t) \tag{4}
$$
As the integration of the differential equation can depend on underlying floating-point arithmetic and thus produce varying time series on different machines, the datasets are precomputed and loaded for training and evaluation. In the benchmark results, 30 instantiations of the Mackey Glass system are used, each with a length of 20 Lyapunov times and successively shifted forwards by half a Lyapunov time. The dataset time series are generated for 50 total Lyapunov times to allow for varied offset starting points. The generated time series are available to be downloaded under the NeuroBench harness.
| 17 | 197 | 0.7206597 |
| --- | --- | --- |
| 18 | 138 | 0.7744313 |
| 19 | 315 | 0.7783468 |
| 20 | 131 | 0.9225991 |
| 21 | 191 | 0.9479431 |
| 22 | 119 | 0.5455960 |
| 23 | 106 | 0.8622247 |
| 24 | 97 | 0.3259660 |
| 25 | 98 | 0.8297825 |
| 26 | 104 | 1.0033490 |
| 27 | 112 | 0.6491406 |
| 28 | 119 | 1.0957495 |
| 29 | 131 | 0.9256179 |
| 30 | 139 | 0.2713639 |
Table 7: Mackey-Glass parameters used for the 14 time series.
Symmetric mean absolute percentage error (sMAPE, Equation 5), a standard metric in forecasting [32], is used to measure the correctness of the model predictions $\hat{y_i}$ against the ground-truth $y_i$ , over $n$ data points in the test split of the time series. The sMAPE metric has a bounded range of $[0,200]$ , thus diverging predictions (infinity or NaN) due to floating-point arithmetic have bounded error which can be used to average correctness over multiple time series instantiations.
$$
sMAPE=200×\frac{1}{n}≤ft(∑_i=1^n\frac{|y_i-\hat{y_i}|}{(|y_
i|+|\hat{y_i}|)}\right) \tag{5}
$$
### Algorithm Track Baselines
All baselines are implemented using PyTorch nn.Module objects in order to interface with the harness.
#### Keyword FSCIL
The ANN baseline employs Mel-frequency cepstral coefficients (MFCC) pre-processing along with a modified version of the M5 deep convolutional network architecture [44].
The MFCC pre-processing converts the 48 kHz, 1 second audio samples from MSWC into 20 channels of 200 timesteps (5 ms stride, 10 ms time bins), focusing on frequencies within the human voice range between 20 Hz and 40 kHz. The network contains four successive blocks, each consisting of 1D convolution, batch-normalization, ReLU activation, and max-pooling layers, followed by a single readout fully-connected layer. Convolutional layers apply their kernels over the temporal dimension of the samples, thus extracting longer temporal features through the depth of the network. We also incorporate dropout after the ReLU activations to avoid over-fitting and let the network be more general for incremental learning. The network is trained with stochastic gradient descent using cross-entropy loss and the Adam optimizer.
For the SNN baseline, we employ the Speech2Spikes [46] (S2S) preprocessing algorithm to convert audio samples to spikes. For S2S we use the default parameters from the original implementation, only the hop length is updated to match the 48 kHz audio frequency of the MSWC samples, whereas the original implementation was applied to 16 kHz audio. S2S applies a Mel Spectrogram and a log operation to raw audio samples, converting them to positive and negative trains of spikes using delta-encoding.
Spike trains from S2S are used as input for the recurrent SNN (RSNN), which consists of 2 recurrent adaptive leaky integrate-and-fire (RadLIF) layers of 1024 neurons and one linear output layer. The model architecture is adapted from Bittar’s work [45]. The RadLIF neurons in these layers are LIF neurons that produce a binary spike $s(t)$ and reset via subtraction when their membrane potential $u(t)$ crosses a certain threshold value $θ$ , combined with an extra adaptation variable $w(t)$ to enable more complex temporal dynamics and firing patterns. Equation 6 is the input current to neurons, with $x$ the input spikes from the previous layer, $W_f$ the forward weight matrix, $BNTT$ batch-normalization through time, and $W_r$ the recurrent weight matrix. $u(t)$ and $w(t)$ are shown in Equation 7, where $α$ , $β$ , $a$ and $b$ are heterogeneously trainable parameters of the neuron. Finally, spikes $s(t)$ are generated according to Equation 8.
$$
I(t)=BNTT(W_f[x(t)])+W_r[s(t-1)] \tag{6}
$$
$$
\displaystyleu(t) \displaystyle=α≤ft[u(t-1)\right]+(1-α)≤ft[I(t)
-w(t-1)\right]-θ[s(t-1)] \displaystylew(t) \displaystyle=β[w(t-1)]+a(1-β)[u(t-1)]+b[s(
t-1)] \tag{7}
$$
$$
s(t)=\begin{cases}0 if u(t)<θ\\
1 if u(t)≥θ\end{cases} \tag{8}
$$
The last layer of the network is a readout linear classifier, and the class corresponding to the maximum of the summation of output activities over all timesteps is chosen as the network prediction. The RSNN network is trained with backpropagation through time using a boxed pseudo-gradient and cross-entropy loss.
Algorithm 1 Few-Shot Class-Incremental Learning with Prototypes
Requires: Pre-trained network $g∘ f$ consisting of feature extractor $f$ and classifier $g:x↦ Wx+b$ Define: $(x)_l$ , $w_l$ and $b_l$ respectively the set of input samples, classifier weights and biases associated with a class $l$
1: for each base class $k$ do
2: Compute prototype embedding $c_k=Mean[f((x)_k)]$ (also summed over time for SNN baseline)
3: Compute corresponding classifier weights $w_k=2c_k$ and biases $b_k=-c_kc_k^T$
4: end for
5: Replace classifier layer $g$ with prototype weights: $W← W_B=(w_k)_k∈ B$ and biases $b← b_B=(b_k)_k∈ B$
6: for each session $i$ in sessions do
7: Get session support $S^i$
8: Repeat lines 1 to 4 for all new classes of $S^i$ to get prototype weights $W_S^i$ and biases $b_S^i$
9: Extend the classifier layer weights $W←[W,W_S^i]$ and $b←[b,b_S^i]$
10: end for
We implement baseline solutions for the FSCIL task with both ANN and SNN models. The frozen baselines do not learn any new classes while the prototypical baselines follow the prototypical networks approach [47] to classify new classes. For both baselines, the ANN and SNN models are pre-trained on the 100 base classes $B$ , which employs the abundant number of samples to develop a robust feature extractor $f$ , which generates embeddings from hidden layers that are passed to a readout classifier.
For the frozen baselines, the models parameters are frozen after pre-training for inference during all incremental sessions, thus setting a ‘worst-case’ reference with no incremental learning but also no risk of catastrophic forgetting.
For the prototypical baselines, the pre-trained models learn 100 extra classes within the 10 incremental sessions in a 5-shot learning scenario. The prototypical networks protocol is applied in each incremental session as shown in Algorithm 1. Prototypical networks provide a clustering algorithm for classification that is equivalent to a readout affine operation on feature embeddings, resulting in a linear layer of weights and biases. Each class $k$ is represented by a prototype vector $c_k=Mean[f((x)_k)]$ defined as the average feature embedding produced by $f$ over all corresponding training samples $(x)_k$ . The readout classifier layer is defined based on this prototype such that the weights $w_k$ and biases $b_k$ associated with class $k$ follow $w_k=2c_k$ and $b_k=-c_kc_k^T$ , which associates embeddings with the closest prototype with respect to the squared Euclidean distance [47].
For the SNN baselines, as the features also have a temporal dimensionality, we accumulate embeddings over all timesteps $t$ to define the prototype vector $c_k=Mean[∑_t(f((x)_k)_t)]$ . Also, as we maintain the summation over timesteps after the final prototype layer to keep the online nature of the SNN baseline, the biases will be applied at each timestep. Thus to maintain the balance between weighted inputs and biases, for the SNN baseline we also normalize the biases by the total number of timesteps $T$ : $b_k=-c_kc_k^T/T$ .
We fit the prototypical networks approach to the FSCIL task by first discarding the original output layer and replacing it with the prototype weights $W_B$ and biases $b_B$ of the base classes, computed as described above based on the averaged feature embeddings over all 500 training samples per base class. This causes an initial accuracy drop, as the trained output layer weights are replaced by clustered weights for the prototypical learning approach. Then, for each incremental session, the prototype of each of the 10 new classes is defined based on the 5 corresponding support samples. The prototype weights and biases are computed in the same manner and concatenated to the existing classifier layer to accommodate for the new classes.
#### Event Camera Object Detection
For both the RED ANN and Hybrid ANN-SNN baselines, the event data from the event camera are converted into frame-based representations using multi-channel time surfaces. Non-overlapping 50 ms time bins (with 50 ms stride), are further subdivided into three sub-bins. Each sub-bin, starting at timestamp $t_0$ , generates two time surfaces $TS$ (Equation 9), based on each event $(x,y,p,t)$ in the sub-bin, where $x,y$ are event coordinates, $p$ is positive or negative polarity, and $t$ is the event time.
$$
\displaystyle TS(p,y,x)=t-t_0 for each event (x,y,p,t) in the
sub-bin. \tag{9}
$$
The RED ANN [22] is a deep convolutional neural network model using three feed-forward squeeze-and-excite [48] convolution layers followed by five recurrent convolution-LSTM [49] (ConvLSTM) layers. The squeeze-and-excite layers provide effective feature extraction while the ConvLSTM layers provide effective temporal learning. The single-shot detection (SSD [50]) head is used to predict the location and class of the bounding box based on multi-scale outputs from the recurrent layers.
The Hybrid ANN-SNN architecture adopts five LIF spiking neural layers to replace the ConvLSTM layers in RED, and shares the same feed-forward convolutional blocks as the RED. The LIF neuron layers are connected with feed-forward convolution, and have far fewer weights than the ConvLSTM layers. The Hybrid model uses the same input encoding method, object detection head, and training loss functions as the RED model. The LIF units are built using the SpikingJelly library [35], and the neuron dynamics of the LIF membrane potential are given in Equations 10, 11, and 12. $h(t)$ is the charged potential before spiking during a timestep, dependent on activation input $X(t)$ , and membrane time contant $τ$ , and $u(t)$ is the final potential of the timestep which resets to the reset value $V_reset$ if $h(t)$ reaches the threshold voltage $V_th$ . The same thresholds determine $s(t)$ , whether a spike is produced. In the experiments, $τ$ is set to 2.0; $V_th$ is 1.0, and $V_reset$ is 0.0.
$$
h(t)=u(t-1)+\frac{1}{τ}(X(t)-u(t-1)) \tag{10}
$$
$$
u(t)=\begin{cases}h(t)&if h(t)<V_th\\
V_reset&if h(t)≥ V_th\end{cases} \tag{11}
$$
$$
s(t)=\begin{cases}0&if h(t)<V_th\\
1&if h(t)≥ V_th\end{cases} \tag{12}
$$
The losses used to train the RED ANN and Hybrid baselines match previous work [22], using a combination of regression and classification loss functions. Regression loss $L_r$ (Equation 13) for all predicted boxes $B$ and ground-truth boxes $T$ is given by smooth $l1$ loss $L_s$ [50] (Equation 14), averaged over $N$ predicted bounding boxes $B_i$ and their corresponding ground-truth boxes $T_i$ . Smooth $l1$ loss is a piecewise loss function with threshold $β$ , which is set to 0.11. For the classification loss ${L}_c$ (Equation 15), softmax focal loss [83] is used, with correct-class probability $p_l$ for all default boxes in the regression head and constant $γ$ , which is set to 2.
$$
\displaystyle L_r(B,T)=\frac{1}{N}∑_jL_s≤ft(B_i,T_i\right) \tag{13}
$$
$$
{L}_s≤ft(B_i,T_i\right)=\begin{cases}≤ft|B_i-T_i\right|-\frac{
β}{2}& if ≤ft|B_i-T_i\right|≥β\\
\frac{1}{2β}≤ft(B_i-T_i\right)^2& otherwise \end{cases} \tag{14}
$$
$$
\displaystyle L_c≤ft(p_{l}\right)=-≤ft(1-p_{l}\right)^γ\log
≤ft(p_{l}\right) \tag{15}
$$
#### Non-human Primate Motor Prediction
All baseline models have linear feed-forward layer architectures, where ANN, ANN_Flat, and SNN_Flat have topologies $N_ch-32-48-2$ , and SNN uses $N_ch-50-2$ . The varying topologies between SNN and SNN_Flat attempt to optimize for complexity in the former and correctness in the latter.
The LIF neurons used in the SNN networks are developed using snnTorch [18], and have potential dynamics shown in Equations 16 and 17. Note that unlike the SpikingJelly neurons (Equations 10, 11, and 12), the potential $u(t)$ is reset in the timestep following a spike, rather than during the same timestep. As before, $V_reset$ is 0.0 and $V_th$ is 1.0, while $β$ is 0.96 for the SNN baseline and 0.50 for the SNN_Flat baseline. The potential of the readout neurons in both baselines is directly read to produce velocity predictions, thus there is no spiking or reset mechanism and the neurons function as leaky accumulators.
$$
u(t)=\begin{cases}βu(t-1)+X(t)&if s(t-1)
=0\\
V_reset&if s(t-1)=1\end{cases} \tag{16}
$$
$$
s(t)=\begin{cases}0&if u(t)≤ V_th\\
1&if u(t)>V_th\end{cases} \tag{17}
$$
ANN, ANN_Flat, and SNN_Flat are trained using mean-squared error (MSE) loss over 50 epochs. The SNN baseline used a sliding window of 50 consecutive data points, representing 200 ms of data (50-point window, single-point stride) in order to calculate the loss, to allow for more information for backpropagation and avoid dead neurons and vanishing gradients. The MSE loss was linearly weighted from 0 to 1 for the 50 points within the window. The SNN was trained with 10-fold cross-validation, using an early-stopping regime with patience (epochs for which there is no improvement to the validation set) of 10 epochs.
#### Chaotic Function Prediction
The LSTM baseline uses one LSTM layer followed by a ReLU activation and linear readout layer. As input, the LSTM uses an explicit memory buffer of the last $M=50$ points. During training, input $x(t)$ to the LSTM uses the Mackey-Glass data $f(t)$ (Equation 18), whereas, during autoregressive evaluation, the input uses prior predictions $y(t)$ (Equation 19). Values $u(t<0)$ and $v(t<0)$ are zero.
$$
x(t)=(f(t-M),f(t-M-1),…,f(t)) \tag{18}
$$
$$
x(t)=(y(t-M-1),y(t-M-2),…,y(t-1)) \tag{19}
$$
The LSTM is trained using MSE loss for backpropagation with 200 epochs. The hyperparameter sweep used the evaluation setup of 30 instantiations of $τ=17$ Mackey-Glass data, with each instance shifted forward by half of the Lyapunov time. The corresponding sets with the lowest sMAPE scores were used to report the results.
For the ESN, the standard architecture with one hidden layer (i.e., reservoir) with recurrent connections was used, where the states of the reservoir $r(t)∈ℝ^D$ at timesteps $t$ are evolving according to the dynamics shown in Equation 20. The random matrix $W^in∈ℝ^D× d+1$ with components drawn from the uniform distribution projects $d$ -dimensional input $f(t)$ ( $d=1$ for the Mackey-Glass system), augmented with constant bias, into $D$ neurons of the reservoir. The recurrent connectivity is defined by the second (potentially sparse) random matrix $W∈ℝ^D× D$ with nonzero components drawn from the normal distribution; $α$ , $γ$ , and $β$ are hyperparameters controlling the behavior of the ESN.
$$
r(t)=(1-α)r(t-1)+α\tanh≤ft(γW
r(t-1)+βW^in[1;f(t)]\right) \tag{20}
$$
To make a prediction $y(t)$ , the ESN uses the readout matrix $W^out∈ℝ^d× D+d+1$ that computes the activation of the output layer based on the current states of the input and hidden layers: $y(t)=W^out[f(t);r(t)]$ . To predict the values of the system at the next timestep, i.e. $y(t)$ predicts $f(t+1)$ , the output layer has $d$ neurons.
The training of $W^out$ is formulated as a linear regression problem so that it can be computed with the regularized least squares estimator (Equation 21), where $H∈ℝ^M× D+d+1$ is an activation matrix that stores the readout for $M$ timesteps in the training data, $Y∈ℝ^M× d$ is another matrix that stores the corresponding ground-truth values for the same timesteps, and $λ$ is the regularization parameter of the estimator.
$$
W^out=Y^⊤H≤ft(H^⊤
H+λI\right)^-1 \tag{21}
$$
Like for the LSTM, optimal hyperparameters are chosen based on lowest average sMAPE score over 30 time series. For each series, the ESN weight matrices $W^in$ and $W$ were randomly initialized. The corresponding sets with the lowest sMAPE scores were used to report the results.
### System Track Metrics
Given the variability of task application areas and system sizes, the NeuroBench metric methodology is defined on a per-benchmark basis. Particularly for efficiency metrics, as neuromorphic systems are not matured, a singular power measurement method cannot be applied to all submissions. It is the responsibility of the submitter to faithfully capture all active processing components for their system and fairly outline their methodology in the report. As NeuroBench moves towards head-to-head benchmarking of neuromorphic systems by hardware vendors and owners, official results must be associated with a report that provides context for the benchmark submission results, as the overall benchmark format is generally open and does not have stringent consistency rules. The report must contain
- an outline of the system architecture,
- an outline of the algorithm used (model architecture, tuning),
- a diagram depicting the workflow, including (where applicable) data initialization, host pre-processing, data loading, on-device preprocessing, inference, post-processing,
- timing measurement description,
- power measurement description, including measurement devices, included hardware components, and measurement time resolution,
- a re-iteration of the results.
In addition, official submissions will be subject to potential audits during which an auditor will inspect the methodology and request additions or revisions to the results and report if necessary.
### System Track Benchmarks
Detailed information on each v1.0 system track benchmark is provided here, and the most updated information can be found in the official NeuroBench system track documentation: https://github.com/NeuroBench/system_benchmarks
#### Acoustic Scene Classification (ASC)
Benchmark Dataset
The dataset is based on the DCASE 2020 acoustic scene classification challenge [68], using the TAU Urban Acoustic Scenes 2020 Mobile datasets. Of the 10 available scene classes, 4 are used: “airport”, “street_traffic”, “bus”, “park”. Of the 9 real / simulated audio recoding devices available, 1 real device is used: “a”. Audio samples are sliced into 1-second samples. The audio may be resampled to a different frequency as a pre-processing step, which is not included in inference measurement. 41360 training samples are available, as well as 16240 test samples. The NeuroBench system track repository linked above provides a download script and a PyTorch-compatible dataset file which is expected to be used as the front-end data generator for all submissions. The data may be reformatted into a different framework, this is not included in inference measurement.
Task
After training a model using the training set, the submitter will report test set accuracy, as well as execution time and energy per inference on the system. The test split should not be used to train or tune the model, only the train split. Audio samples will be processed in batch-size 1, in which one sample is processed at a time and the next sample is not processed until the previous one is finished. The general compute flow consists of three steps: pre-processing, inference, and classification. One inference is defined as the processing of one second of audio data. Inference is separate from classification because systems are often intended to process samples as sequences over time, rather than all at once, and the classification may be available before the whole data sequence is seen. Classification thus does not need to be included in the benchmark measurement. Pre-processing in general can be defined as feature extraction or the conversion of raw audio data into a format that is suitable for inference. Execution time and energy of pre-processing must be included in benchmark measurement, and will be reported separately from inference. In certain systems, e.g. Synsense Xylo, pre-processing blocks use analog hardware on-chip to directly convert real-time analog microphone output to digital spikes for inference. In order to operate inference from a digitally-encoded dataset, this pre-processing must be simulated and the spikes are sent through a side-channel directly to the inference block. For such cases, the reported pre-processing measurements will be for the analog hardware running in real-time on the dataset audio played over speakers into the device microphone.
Metrics
The following metrics should be reported:
- Accuracy – Accuracy of the predictions on the test set, measured from the system and not in any software simulation.
- Execution time – Execution time is the average time per sample for pre-processing and inference. The final result should be averaged over all samples in the test set. The time begins when data has been loaded into on-board or on-chip memory and is prepared to be processed. The time ends when the last timestep of the sequence has completed processing.
- Power/Energy – Idle power and active power should be reported, where idle power is power of the chip when it has been configured and it has prepared to begin processing. Note that this should not measure the device in a lower-power sleeping state. Active power is the average power consumed during processing. The difference between active power and idle power should be reported as dynamic power. Power should be converted to energy by multiplying power by the averaged execution time. The power measurement should include all computational components and power domains which are active during the workload processing. If applicable, power may be measured over longer data sequences (e.g. 5 seconds rather than 1 second), such that configuration costs are amortized over a longer period of processing. This should be done separately from accuracy and execution time experiments, and should be clearly detailed in the associated submission report.
#### Quadratic Unconstrained Binary Optimization (QUBO)
Quadratic Unconstrained Binary Optimization (QUBO) refers to the problem of finding the binary variable assignment $x_i∈\{0,1\}$ that optimizes the quadratic cost function
$$
\min_x∈\{0,1\^n}c(x)=\min_x∈\{0,1\^n}
x^TQx
$$
subject to no constraints.
The solvers for QUBO will be benchmarked using Maximum Independent Set (MIS) workloads. Given an undirected graph $G=(V,E)$ , an independent set $I$ is a subset of $V$ such that, for any two vertices $u,v∈I$ , there is no edge connecting them, i.e., $≠xists e∈E s.t. e=(u,v) \vee e=(v,u)$ .
The MIS problem has a natural QUBO formulation: for each node $u∈V$ in the graph, a binary variable $x_u$ is introduced to model the inclusion or not of $u$ in the candidate solution. Summing the quadratic terms $x_u^2$ will thus result in the size of the set of selected nodes. To penalize the selection of nodes that are not mutually independent, a penalization term is associated to the interactions $x_ux_v$ if $u$ and $v$ are connected. The resulting $Q$ matrix coefficients are defined as
$$
q_uv=\begin{cases}-1&if u=v\\
4&if u≠ v and (u,v)∈E\\
0&otherwise.\end{cases}
$$
The MIS problem is NP-hard and intractable, and therefore any solver system approximates a solution. Therefore, the cost of the QUBO formulation ( $x^TQx$ ) is used to assess solutions, and solutions are not restricted to being maximum nor independent sets. The QUBO formulation ensures that any non-independent set will always have a higher cost than a corresponding independent set with conflicting nodes removed.
Benchmark Dataset
The benchmark’s workload complexity is defined such that it can automatically grow over time, as neuromorphic systems mature and are able to support larger problems.
- Number of nodes, spaced in a pseudo-geometric progression: 10, 25, 50, 100, 250, 500, 1000, 2500, 5000, …
- Density of edges: 1%, 5%, 10%, 25%.
- Problem seeds: 0, 1, 2, 3, 4 are allowed for tuning. At evaluation time, for official results NeuroBench will announce five seeds for submission. Unofficial results may use seeds which are randomly generated at runtime.
Each workload will be associated with a target optimality, which is the minimum cost found using a conventional solver algorithm. Small QUBO workloads with fewer than 50 nodes will be solved to global optimality, corresponding to the true maximum independent set. Larger workloads cannot be reasonably globally solved. The DWave Tabu CPU sampler [74] will be used with 100 reads and 50 restarts, and the QUBO solution with the best cost (best-known solution, BKS) found will set the target optimality for the tuning workload seeds. For evaluation workload seeds, the same method will be used to set the target optimality. NeuroBench will provide target optimalities for workloads up to 5000 nodes. Submissions are encouraged to continue scaling up the workload size along the pattern to demonstrate the capacity of their systems. The first group that tackles workloads of an unprecedented size should provide the benchmark solutions via a pull request to the system track repository. The dataset workload generator and scripts for the DWave Tabu sampler to compute optimal costs are available in the repository, and this code is expected to be used as the front-end data generator for all submissions.
Task and Metrics
Based on the BKS found for each workload, the BKS-Gap optimality score of the solution found by the SUT is defined as
$$
BKS-Gap=(\frac{c-c_target}{c_target}) ,
$$
where $c_target$ is the QUBO cost of the BKS, and $c$ is the cost found by the SUT. This may be reported as a percentage gap by multiplying by 100. If the SUT manages to beat the BKS, then the BKS-Gap will be negative.
Given each workload, the benchmark should report the BKS-Gap of the solution found by the SUT after a fixed runtime. The time begins after the graph has been loaded into the SUT. Timeouts spread across orders of magnitude ( $10^-3$ , $10^-2$ , $10^-1$ , $10^0$ , $10^1$ , $10^2$ ). As the runtime is fixed, no measured timing metric is reported, and submissions should report average power over the duration of the runtime, as it is directly proportional to energy consumed. Importantly, the QUBO solver needs a module to measure the cost of its solutions, and this module should be considered as part of the SUT, thus its computational demand and power must be included in the benchmarking results.
In the future, a different optimization task scenario may be used for the same QUBO dataset, in which the SUT must run until it reaches a small BKS-Gap, rather than stopping after a fixed-timeout. Here, the benchmark should report latency and energy required to reach BKS-Gap thresholds, e.g., 0.1, 0.05, and 0.01. This task scenario normalizes systems to solution quality, rather than SUT runtime.
### System Track Baselines
#### Acoustic Scene Classification (ASC)
<details>
<summary>x9.png Details</summary>

### Visual Description
## Diagram: System Architecture for Xylo HDK
### Overview
This diagram illustrates the system architecture and data flow for the Xylo Hardware Development Kit (HDK). It depicts a dual-mode operation environment: a simulation mode where a PC provides data to the Xylo hardware, and a live mode where a physical microphone provides input. The system is designed for testing and deploying Spiking Neural Network (SNN) applications.
### Components/Axes
The diagram is divided into two primary regions: the **PC (Left)** and the **Xylo HDK (Right)**.
**Left Region (PC):**
* **PC Icon:** Located at the top-left.
* **Dataset:** A box containing a hard drive icon and the label "Dataset".
* **AFESim:** A box labeled "AFESim" with the sub-label "Audio encoding" positioned directly below it.
* **Result collection:** A box labeled "Result collection" located below the AFESim box.
**Right Region (Xylo HDK):**
* **Container:** A large, rounded green rectangle labeled "Xylo HDK" at the top center.
* **Mic.:** A box containing a circle icon and the label "Mic." located in the top-left of the green container.
* **Xylo™ Audio 2:** A sub-box located in the top-right of the green container.
* **AFE:** A box inside the Xylo™ Audio 2 container with the sub-label "Audio encoding" below it.
* **SNN core:** A box inside the Xylo™ Audio 2 container with the sub-label "Inference" below it.
* **FPGA:** A box located in the bottom-left of the green container.
* **Configuration Power measurement:** A label positioned to the right of the FPGA box.
### Connections
* **USB:** A label placed above the line connecting the PC region to the Xylo HDK region.
* **Data Flow Lines:**
* Solid line from "Dataset" to "AFESim".
* Solid line from "AFESim" to "SNN core" (labeled "USB" at the connection point).
* Dashed line from "Mic." to "AFE".
* Solid line from "AFE" to "SNN core".
* Solid line from "SNN core" to "Result collection".
### Detailed Analysis
The diagram outlines two distinct workflows for the "SNN core":
1. **Simulation/Development Path:**
* Data originates from the "Dataset" on the PC.
* It passes through "AFESim" (Analog Front End Simulation), which performs "Audio encoding".
* The encoded data is sent via "USB" to the "SNN core" on the Xylo™ Audio 2 chip.
* The "SNN core" performs "Inference" and sends the output back to the "Result collection" block on the PC.
2. **Live Hardware Path:**
* Audio is captured by the "Mic." (Microphone).
* The signal is sent to the "AFE" (Analog Front End) on the Xylo™ Audio 2 chip.
* The AFE performs "Audio encoding" and passes the signal to the "SNN core".
* The "SNN core" performs "Inference".
**Supporting Components:**
* The **FPGA** is present within the Xylo HDK environment, likely serving as the controller for system configuration and power monitoring, as indicated by the label "Configuration Power measurement".
### Key Observations
* **Dual-Input Architecture:** The "SNN core" is the central processing hub, capable of receiving inputs from two distinct sources: the PC (via USB/AFESim) or the physical Microphone (via AFE).
* **Functional Parity:** The presence of "Audio encoding" labels under both "AFESim" and "AFE" suggests that the simulation environment is designed to mirror the hardware's analog front-end processing exactly.
* **Feedback Loop:** The system is closed-loop for the PC, as it both provides the input (Dataset) and collects the output (Result collection).
### Interpretation
This diagram represents a typical "Hardware-in-the-loop" (HIL) or development environment for neuromorphic computing.
* **Why it matters:** By providing an "AFESim" (Simulation) path, developers can validate their SNN models against pre-recorded datasets without needing to physically interact with the microphone or the Xylo hardware in real-time. This is crucial for debugging and training.
* **Peircean Investigative View:** The "Xylo™ Audio 2" is clearly the System-on-Chip (SoC) or primary processing unit. The separation of "AFE" and "SNN core" suggests a modular design where the analog signal processing is distinct from the neural network inference. The inclusion of the FPGA for "Configuration Power measurement" indicates that this is a development board (HDK) intended for power-sensitive edge AI applications, where monitoring energy consumption is as important as the inference accuracy.
</details>
<details>
<summary>x10.png Details</summary>

### Visual Description
## PCB Layout: SynSense XYLO-AUDIO-V2 Daughter Board
### Overview
The image displays a high-resolution, top-down view of a black printed circuit board (PCB) assembly, identified as the "XYLO-AUDIO-V2 DAUGHTER BOARD V2.0". The assembly consists of a main board and a secondary board (likely an interface or bridge board) attached at the top. The board is designed for neuromorphic audio processing, featuring a central processing chip, various power management components, and an extensive array of test points for debugging and signal monitoring.
### Components/Axes
The board is organized into functional regions:
**1. Top Section (Interface/Bridge Board):**
* **FX3 TESTPOINT Block (Top Right):** A cluster of test points labeled:
* T1 UART TXD
* T2 UART CTS
* T3 UART RTS
* T4 UART RXD
* T5 I2S CLK
* T6 I2S SD
* T7 I2S WS
* T8 I2S SCK
* **Miscellaneous:** Various surface-mount capacitors, resistors, and an integrated circuit (U18) are visible.
**2. Main Board (Center/Left):**
* **Branding:** "SynSense" logo accompanied by Chinese characters: **时 | 识 | 科 | 技** (Translation: Time | Knowledge | Science | Technology).
* **Product Logo:** "XYLO" logo (stylized bee/insect icon).
* **Input/Power:**
* "External Signal In" and "MICIn" (Microphone Input) near a gold SMA connector.
* Power rail indicators: "POWER", "IO_POWER", "CORE_POWER".
* **Central Processor:** A square integrated circuit located in the center of the board, highlighted by a red square border.
* **Identification Label:** A white sticker reading "SYN61202_04".
**3. Main Board (Right - Pinout/Test Points):**
A vertical column of 24 test points/pins, labeled from top to bottom:
* SPI_SSN1
* SPI_SSN0
* SPI_SCLK
* SPI_MOSI
* SPI_MISO
* MON_05
* MON_03
* MON_04
* MON_02
* MON_01
* MON_00
* INT_0
* TDO
* CLKIN
* TDI
* TMS
* TCK
* TRST
* RST_N
* CLK_I
* SAER_DATA_VLD
* SAER_DATA
* SAER_CLK
* GND
**4. Footer:**
* "XYLO-AUDIO-V2 DAUGHTER BOARD V2.0"
* "403988Y-Y19-221017" (Serial/Part number)
### Detailed Analysis
* **Language:** The Chinese characters "时 | 识 | 科 | 技" are present. These translate to "Time | Knowledge | Science | Technology," which aligns with the "SynSense" brand identity (a neuromorphic computing company).
* **Connectivity:** The board is heavily focused on connectivity. The presence of "FX3" test points suggests the board is designed to interface with a Cypress/Infineon FX3 USB 3.0 peripheral controller, which is common for high-speed data streaming from sensors.
* **Neuromorphic Indicators:** The labels "SAER_DATA" (SynSense Asynchronous Event Representation) and "MON_XX" (Monitor pins) are characteristic of neuromorphic hardware, which processes data as asynchronous spikes rather than traditional clocked frames.
* **Power Management:** The board includes dedicated power rails (POWER, IO_POWER, CORE_POWER), suggesting the need for distinct voltage domains for the analog/sensor front-end and the digital processing core.
### Key Observations
* **Red Box:** The red box highlights the primary processing chip (likely the XYLO chip itself).
* **Modularity:** The board is explicitly labeled as a "Daughter Board," indicating it is intended to be plugged into a larger "Motherboard" or evaluation system.
* **Debugging:** The high density of test points (UART, I2S, SPI, JTAG/SWD, and Monitor pins) indicates this is an engineering/development version of the hardware, not a final consumer product.
### Interpretation
This image depicts a development platform for SynSense's neuromorphic audio processing technology.
* **Functionality:** The board is designed to ingest audio signals (via MICIn or External Signal In), process them using the central chip (SYN61202), and output the processed data—likely in an event-based format (SAER)—to an external host via the SPI or I2S interfaces.
* **Peircean Investigative View:** The presence of "FX3" test points and "SAER" pins suggests a high-bandwidth data pipeline. The board is designed to bridge the gap between raw analog audio input and digital event-based processing. The "Daughter Board" designation implies that the user is expected to provide the power and host-interface logic (likely via the FX3 controller) on a separate carrier board, allowing the user to swap out different sensor or processing modules easily. The board is a tool for rapid prototyping and algorithm validation in the field of neuromorphic engineering.
</details>
Figure 9: Left: An overview of the benchmarking system for the Xylo. A simulator of the analog encoding module is run on a PC and streamed to the SNN inference core via USB. After inference, outputs are routed back to the PC for classification. An on-board FPGA configures and records power of the Xylo components. Right: The Xylo™Audio 2 hardware development kit (HDK), used as the SUT. The red outline marks the Xylo inference module. Figures taken, with permission, from Ke et al. [72]
CPU
The CPU baseline uses an Arduino Nano 33 BLE board, which runs preprocessing and inference on an ARM Cortex-M4F microcontroller running at 64MHz. All training and deployment uses Edge Impulse, a commercially developed ML-operations platform for tinyML [84]. The 1 second digital audio samples were converted into two-dimensional frames of Mel-filterbank energies (MFE), and inference uses a network with two convolution layers with batch normalization and max pooling. The trained model was quantized then compiled to an Arduino library using Edge Impulse EON compiler, which optimizes for memory and flash usage. Execution time is measured using on-chip timers from the Arduino, while power is measured using an external multimeter for total system power. Power was measured separately for idle, active pre-processing, and active inference by taking the average power over 60 seconds. The Arduino repeatedly computes over one sample loaded in memory, with a recording frequency of 3 Hz on a Keysight 34465A digital multimeter.
Synsense Xylo
As shown in Figure 9, the Xylo baseline used a host PC to run a simulation of the analog pre-processing unit on the benchmark dataset, which was routed into the SNN core on the Xylo SUT. Accuracy is measured by performing a maximum on the output spikes of the SNN. The SNN on Xylo is a feedforward network with three hidden layers, where each layer is connected by synapses with varying time delays. Training is done using the open-source Rockpool toolchain [85]. In order to amortize measurement overheads, power and execution time are measured over continuous streams of 10 seconds of audio, where the execution time result is divided by 10 and the power result is averaged over the duration. The measurements are made by the on-board FPGA, which operates at 12.5MHz and samples power at 1280 Hz. Accuracy is still measured on the 1-second samples. Full details are available in Ke et al. [72]
#### Quadratic Unconstrained Binary Optimization (QUBO)
<details>
<summary>x11.png Details</summary>

### Visual Description
## Diagram: System Architecture for QUBO Workload Processing on Loihi 2
### Overview
The image illustrates a hardware-in-the-loop system architecture where a host computer (PC) offloads a specific computational task—a "QUBO workload"—to an Intel "Kapoho Point board" equipped with Loihi 2 neuromorphic chips. The board processes the workload and returns a "Solution & its cost." The right side of the image provides a schematic representation of the internal architecture of the Loihi 2 chip, highlighting its grid-based core structure and I/O interfaces.
### Components/Axes
**1. Left Section (Host System):**
* **PC Icon:** Represents the host computer.
* **Ethernet Icon:** Represents the network connection (switch/hub symbol) between the PC and the hardware board.
* **Flow Arrows:**
* **Top Arrow:** Labeled "QUBO workload" (orange text), pointing from the PC/Ethernet system toward the Kapoho Point board.
* **Bottom Arrow:** Labeled "Solution & its cost" (orange text), pointing from the Kapoho Point board back toward the PC/Ethernet system.
**2. Center Section (Hardware):**
* **Kapoho Point board:** A physical circuit board featuring multiple integrated circuit chips.
* **Highlight:** One specific chip on the board is outlined with a red square, which visually connects to the Loihi 2 schematic on the right.
**3. Right Section (Schematic):**
* **Loihi 2:** A schematic diagram of the chip architecture, enclosed in a red border.
* **Grid Structure:** A matrix of small squares representing neuromorphic processing cores.
* **Parallel IO:** Four blocks labeled "Parallel IO" located at the top, bottom, left, and right edges of the chip schematic.
* **FPIO:** A block labeled "FPIO" (Flexible Pin IO) located at the bottom center of the schematic.
* **Highlight:** A vertical column of cores in the right-center of the grid is highlighted in blue.
### Detailed Analysis
* **Workflow:** The system operates as a request-response loop. The PC transmits a Quadratic Unconstrained Binary Optimization (QUBO) problem to the Kapoho Point board. The board, utilizing the Loihi 2 architecture, computes the optimization and returns the resulting solution along with the associated cost (the value of the objective function).
* **Hardware Architecture (Loihi 2):**
* The schematic depicts a highly parallel, tiled architecture.
* The "Parallel IO" blocks suggest high-bandwidth connectivity for data ingestion and extraction.
* The "FPIO" (Flexible Pin IO) indicates configurable input/output pins, likely used for interfacing with external sensors or other hardware components.
* The blue-highlighted column of cores suggests that the QUBO workload may be mapped to a specific subset of the chip's neuromorphic fabric, or that this region is currently active/being monitored.
### Key Observations
* **QUBO Focus:** The explicit mention of "QUBO" (Quadratic Unconstrained Binary Optimization) indicates that this specific hardware setup is being utilized for combinatorial optimization tasks, which is a non-traditional use case for neuromorphic chips (typically used for spiking neural networks).
* **Spatial Mapping:** The red highlight on the physical Kapoho Point board directly corresponds to the red-bordered Loihi 2 schematic, confirming that the schematic represents the internal logic of the chips mounted on the board.
* **Closed-Loop Optimization:** The return of "cost" alongside the "solution" is critical; it implies the system is performing iterative optimization where the chip minimizes an energy function to find the best solution.
### Interpretation
This diagram demonstrates the repurposing of Intel's neuromorphic Loihi 2 hardware for solving optimization problems.
* **Why it matters:** Traditional CPUs often struggle with the massive parallelism required for certain combinatorial optimization problems. By mapping QUBO problems onto the Loihi 2's neuromorphic fabric, the system can potentially achieve significant speedups or energy efficiency gains.
* **Reading between the lines:** The blue column in the Loihi 2 schematic is likely not arbitrary; it represents a specific mapping of the QUBO problem onto the chip's grid. The "FPIO" and "Parallel IO" labels suggest that the chip is designed to be integrated into larger systems where real-time data throughput is required, rather than just being a standalone processor. The "cost" return is the most significant indicator of the chip's function here: it is acting as an energy-minimization engine, where the "cost" corresponds to the energy state of the neural network representing the QUBO problem.
</details>
Figure 10: A Kapoho Point board with Loihi 2 chips is connected to a host PC via ethernet. Shown here is a Kapoho Point board with 8 Loihi 2 chips, while the experiments were run on a board with 1 Loihi 2 chip. The QUBO workload is loaded onto the chip, and all processing happens on Loihi 2. On Loihi 2, the neuro-cores, in yellow, are arranged in two $8×{}8$ grids and iteratively update the variables of the QUBO workload. Embedded CPUs, shown in blue, monitor the cost of the variable assignment. Six parallel IO ports enable 3D stacking of multiple chips for larger workloads than used here. An additional IO port provides the communication to the host PC. The energy and runtime measurements cover the computation of the Loihi 2 chip.
The processors ran repeatedly with five different initial variable assignments and seeds for different timeouts between $10^-3$ s to $10^3$ s, and for different workloads sizes of up to $1000$ variables. The neuromorphic algorithm is a parallelized version of simulated annealing that was running on a Kapoho Point board with one Loihi 2 chip, as shown in Figure 10. The Loihi 2 board was controlled using Lava 0.8.0 and Lava Optimization 0.3.0. For comparison with conventional hardware, two solvers were adopted based on simulated annealing and tabu search, as implemented in the D-Wave Samplers v1.1.0 library [74]. The library was compiled on Ubuntu 20.04.6 LTS with GCC 9.4.0 and Python 3.8.10. CPU measurements were obtained on a machine with Intel Core i9-7920X CPU @ 2.90 GHz and 128GB of DDR4 RAM, using Intel SoC Watch for Linux OS 2023.2.0. All details on the solvers, benchmarking routine, and results have been provided by Pierro et al. [57].
## Acknowledgements
Authors of this work have been supported in parts by Semiconductor Research Corporation (JY), the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (grant agreement No. 101001448), a grant from the Research Grants Council of the Hong Kong Special Administrative Region, China [Project No. CityU 11200922], ARC Laureate Fellowship FL210100156, and the EU H2020 project BeFerroSynaptic (871737). We acknowledge the financial support of the CogniGron research center and the Ubbo Emmius Funds (Univ. of Groningen). We acknowledge a contribution from the Italian National Recovery and Resilience Plan (NRRP), M4C2, funded by the European Union –NextGenerationEU (Project IR0000011, CUP B51E22000150006, “EBRAINS-Italy”). The work of SynSense was partially supported by the European Commission, under the Horizon grant Ferro4Edge AI (grant agreement 101135656). This work is partly funded by the German Federal Ministry of Education and Research (BMBF) and the free state of Saxony within the ScaDS.AI center of excellence for AI research and by the German Federal Ministry for Economic Affairs and Climate Action (BMWK) under contract 01MN23004F (ESCADE).
Sandia National Laboratories is a multi-mission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC (NTESS), a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration (DOE/NNSA) under contract DE-NA0003525. This written work is authored by an employee of NTESS. The employee, not NTESS, owns the right, title and interest in and to the written work and is responsible for its contents. Any subjective views or opinions that might be expressed in the written work do not necessarily represent the views of the U.S. Government. The publisher acknowledges that the U.S. Government retains a non-exclusive, paid-up, irrevocable, world-wide license to publish or reproduce the published form of this written work or allow others to do so, for U.S. Government purposes. The DOE will provide public access to results of federally sponsored research in accordance with the DOE Public Access Plan. This paper describes objective technical results and analysis. Any subjective views or opinions that might be expressed in the paper do not necessarily represent the views of the U.S. Department of Energy or the United States Government.
## Author contributions statement
Authors are grouped based on contributions, and ordered alphabetically within groups. JY led project discussions and management. JY and KVdB implemented harness metric infrastructure, conducted experiments, and prepared the manuscript. The following authors primarily developed the main algorithm track results: DdB and MF on the keyword few-shot class-incremental learning task; GT and SW on the event camera object detection task; PH, PVS and BZ on the non-human primate motor prediction task; YB and DK on the chaotic function prediction task; and NP led harness infrastructure development. The following authors primarily developed the main system track results: WK and MAK on the acoustic scene classification task; AP and PS on the QUBO task. SHA, GVJ, BL, AM, AKM, GL, and TS developed components of the algorithm track harness infrastructure. ZA, MA, BA, AGA, CB, AB, PB, S. Bohte, S. Buckley, GC, EC, FC, GdC, A. Danielescu, A. Daram, MD, YD, JE, TF, JF, VF, SF, PMF, WG, AG, HAG, GI, SJ, VK, L. Khacef, JCK, L. Kriener, RK, DK, SL, YL, HM, RM, JMM, CM, KM, DRM, EN, TN, FO, AO, PP, JP, MP, C. Pehle, MAP, C. Posch, AR, YS, CJSS, AvS, J. Schemmel, S. Schmidgall, CS, J. Seo, S. Sheik, SBS, MS, AS, KS, MS, TCS, JT, NT, GU, MV, CMV, BV, AY, and FTZ participated in discussions during meetings, prepared sections for the present manuscript and/or its preprint, and reviewed the manuscript. CF and VJR jointly supervised the project, reviewed the manuscript, and analyzed results.
## References
- [1] Sevilla, J. et al. Compute trends across three eras of machine learning. In 2022 International Joint Conference on Neural Networks (IJCNN), 1–8, DOI: https://doi.org/10.1109/IJCNN55064.2022.9891914 (2022).
- [2] Shankar, S. & Reuther, A. Trends in energy estimates for computing in ai/machine learning accelerators, supercomputers, and compute-intensive applications. In 2022 IEEE High Performance Extreme Computing Conference (HPEC), 1–8, DOI: https://doi.org/10.1109/HPEC55821.2022.9926296 (2022).
- [3] Ray, P. P. A review on tinyml: State-of-the-art and prospects. Journal of King Saud University - Computer and Information Sciences 34, 1595–1623, DOI: https://doi.org/10.1016/j.jksuci.2021.11.019 (2022).
- [4] Schuman, C. D. et al. A survey of neuromorphic computing and neural networks in hardware (2017). https://doi.org/10.48550/arXiv.1705.06963.
- [5] James, C. D. et al. A historical survey of algorithms and hardware architectures for neural-inspired and neuromorphic computing applications. Biologically Inspired Cognitive Architectures 19, DOI: https://doi.org/10.1016/j.bica.2016.11.002 (2017).
- [6] Thakur, C. S. et al. Large-scale neuromorphic spiking array processors: A quest to mimic the brain. Frontiers in Neuroscience 12, DOI: https://doi.org/10.3389/fnins.2018.00891 (2018).
- [7] Mead, C. A. Neuromorphic electronic systems. Proceedings of the IEEE 78, 1629–1636, DOI: https://doi.org/10.1109/5.58356 (1990).
- [8] Schuman, C. et al. Opportunities for neuromorphic computing algorithms and applications. Nature Computational Science 2, 10–19, DOI: https://doi.org/10.1038/s43588-021-00184-y (2022).
- [9] Frenkel, C., Bol, D. & Indiveri, G. Bottom-up and top-down approaches for the design of neuromorphic processing systems: Tradeoffs and synergies between natural and artificial intelligence. Proceedings of the IEEE 111, 623–652, DOI: https://doi.org/10.1109/JPROC.2023.3273520 (2023).
- [10] Davies, M. Benchmarks for progress in neuromorphic computing. Nature Machine Intelligence 1, 386388 (2019).
- [11] Orchard, G., Jayawant, A., Cohen, G. K. & Thakor, N. Converting static image datasets to spiking neuromorphic datasets using saccades. Frontiers in Neuroscience 9, DOI: https://doi.org/10.3389/fnins.2015.00437 (2015).
- [12] Amir, A. et al. A low power, fully event-based gesture recognition system. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7243–7252, DOI: https://doi.org/10.1109/CVPR.2017.781 (2017).
- [13] Cramer, B., Stradmann, Y., Schemmel, J. & Zenke, F. The heidelberg spiking data sets for the systematic evaluation of spiking neural networks. IEEE Transactions on Neural Networks and Learning Systems 33, 2744–2757, DOI: https://doi.org/10.1109/tnnls.2020.3044364 (2022).
- [14] Ostrau, C., Klarhorst, C., Thies, M. & Rückert, U. Benchmarking neuromorphic hardware and its energy expenditure. Frontiers in Neuroscience 16, DOI: https://doi.org/10.3389/fnins.2022.873935 (2022).
- [15] Milde, M. B. et al. Neuromorphic engineering needs closed-loop benchmarks. Frontiers in Neuroscience 16, DOI: https://doi.org/10.3389/fnins.2022.813555 (2022).
- [16] Kulkarni, S. R., Parsa, M., Mitchell, J. P. & Schuman, C. D. Benchmarking the performance of neuromorphic and spiking neural network simulators. Neurocomputing 447, 145–160, DOI: https://doi.org/10.1016/j.neucom.2021.03.028 (2021).
- [17] Gewaltig, M.-O. & Diesmann, M. Nest (neural simulation tool). Scholarpedia 2, 1430 (2007).
- [18] Eshraghian, J. K. et al. Training spiking neural networks using lessons from deep learning. Proceedings of the IEEE 111, 1016–1054, DOI: https://doi.org/10.1109/JPROC.2023.3308088 (2023).
- [19] Reddi, V. J. et al. Mlperf inference benchmark. In Proceedings of the ACM/IEEE 47th Annual International Symposium on Computer Architecture, ISCA ’20, 446–459, DOI: https://doi.org/10.1109/ISCA45697.2020.00045 (IEEE Press, 2020).
- [20] Mattson, P. et al. Mlperf training benchmark. Proceedings of Machine Learning and Systems 2, 336–349 (2020).
- [21] Mazumder, M. et al. Multilingual spoken words corpus. In Vanschoren, J. & Yeung, S. (eds.) Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, vol. 1 (Curran, 2021).
- [22] Perot, E., de Tournemire, P., Nitti, D., Masci, J. & Sironi, A. Learning to detect objects with a 1 megapixel event camera. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20 (2020).
- [23] O’Doherty, J. E., Cardoso, M. M. B., Makin, J. G. & Sabes, P. N. Nonhuman primate reaching with multichannel sensorimotor cortex electrophysiology, DOI: https://doi.org/10.5281/zenodo.788569 (2017).
- [24] Mackey, M. C. & Glass, L. Oscillation and chaos in physiological control systems. Science 197, 287–289 (1977).
- [25] Kudithipudi, D. et al. Biological underpinnings for lifelong learning machines. Nature Machine Intelligence 4, 196–210, DOI: https://doi.org/10.1038/s42256-022-00452-0 (2022).
- [26] Tao, X. et al. Few-shot class-incremental learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2020).
- [27] Gallego, G. et al. Event-based vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 154–180, DOI: https://doi.org/10.1109/TPAMI.2020.3008413 (2022).
- [28] Lin, T.-Y. et al. Microsoft COCO: Common objects in context. In Computer Vision – ECCV 2014, 740–755 (2014).
- [29] Jaeger, H. & Haas, H. Harnessing nonlinearity: Predicting chaotic systems and saving energy in wireless communication. science 304, 78–80 (2004).
- [30] Mukhopadhyay, S. & Banerjee, S. Learning dynamical systems in noise using convolutional neural networks. Chaos: An Interdisciplinary Journal of Nonlinear Science 30, 103125 (2020).
- [31] Chilkuri, N. R. & Eliasmith, C. Parallelizing legendre memory unit training. In International Conference on Machine Learning, 1898–1907 (PMLR, 2021).
- [32] Makridakis, S., Spiliotis, E. & Assimakopoulos, V. The m4 competition: 100,000 time series and 61 forecasting methods. International Journal of Forecasting 36, 54–74, DOI: https://doi.org/10.1016/j.ijforecast.2019.04.014 (2020). M4 Competition.
- [33] Gilpin, W. Model scale versus domain knowledge in statistical forecasting of chaotic systems (2023). https://doi.org/10.48550/arXiv.2303.08011.
- [34] Paszke, A. et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, 8024–8035 (Curran Associates, Inc., 2019).
- [35] Fang, W. et al. Spikingjelly. https://github.com/fangwei123456/spikingjelly (2020).
- [36] Intel. Lava software framework. https://github.com/lava-nc/lava (2021).
- [37] Aimone, J. B., Severa, W. & Vineyard, C. M. Composing neural algorithms with fugu. In Proceedings of the International Conference on Neuromorphic Systems, 1–8 (2019).
- [38] Warden, P. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209 (2018).
- [39] Blouw, P., Choo, X., Hunsberger, E. & Eliasmith, C. Benchmarking keyword spotting efficiency on neuromorphic hardware. In Proceedings of the 7th Annual Neuro-Inspired Computational Elements Workshop, NICE ’19, DOI: 10.1145/3320288.3320304 (Association for Computing Machinery, New York, NY, USA, 2019).
- [40] Fang, W. et al. Incorporating learnable membrane time constant to enhance learning of spiking neural networks (2021). 2007.05785.
- [41] Massa, R., Marchisio, A., Martina, M. & Shafique, M. An efficient spiking neural network for recognizing gestures with a dvs camera on the loihi neuromorphic processor (2021). 2006.09985.
- [42] Lemaire, E. et al. An analytical estimation of spiking neural networks energy efficiency. In Neural Information Processing, 574–587, DOI: https://doi.org/10.1007/978-3-031-30105-6_48 (Springer International Publishing, 2023).
- [43] Fra, V. et al. Human activity recognition: suitability of a neuromorphic approach for on-edge aiot applications. Neuromorphic Computing and Engineering 2, DOI: https://doi.org/10.1088/2634-4386/ac4c38 (2022).
- [44] Dai, W., Dai, C., Qu, S., Li, J. & Das, S. Very deep convolutional neural networks for raw waveforms (2016). 1610.00087.
- [45] Bittar, A. & Garner, P. N. A surrogate gradient spiking baseline for speech command recognition. Frontiers in Neuroscience 16, DOI: https://doi.org/10.3389/fnins.2022.865897 (2022).
- [46] Stewart, K. M., Shea, T., Pacik-Nelson, N., Gallo, E. & Danielescu, A. Speech2spikes: Efficient audio encoding pipeline for real-time neuromorphic systems. In Proceedings of the 2023 Annual Neuro-Inspired Computational Elements Conference, NICE ’23, 71–78, DOI: https://doi.org/10.1145/3584954.3584995 (Association for Computing Machinery, New York, NY, USA, 2023).
- [47] Snell, J., Swersky, K. & Zemel, R. Prototypical networks for few-shot learning. Advances in neural information processing systems 30 (2017).
- [48] Hu, J., Shen, L. & Sun, G. Squeeze-and-excitation networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7132–7141, DOI: https://doi.org/10.1109/CVPR.2018.00745 (2018).
- [49] Shi, X. et al. Convolutional lstm network: A machine learning approach for precipitation nowcasting. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS’15, 802–810 (MIT Press, Cambridge, MA, USA, 2015).
- [50] Liu, W. et al. SSD: Single shot MultiBox detector. In Computer Vision – ECCV 2016, 21–37, DOI: https://doi.org/10.1007/978-3-319-46448-0_2 (Springer International Publishing, 2016).
- [51] Willsey, M. et al. Real-time brain-machine interface in non-human primates achieves high-velocity prosthetic finger movements using a shallow feedforward neural network decoder. Nature Communications 13, 6899, DOI: https://doi.org/10.1038/s41467-022-34452-w (2022).
- [52] Hochreiter, S. & Schmidhuber, J. Long short-term memory. Neural Computation 9, 1735–1780, DOI: https://doi.org/10.1162/neco.1997.9.8.1735 (1997).
- [53] Scardapane, S. & Wang, D. Randomness in neural networks: an overview. Data Mining and Knowledge Discovery 7, 1–18, DOI: https://doi.org/10.1002/widm.1200 (2017).
- [54] Bos, H. & Muir, D. R. Micro-power spoken keyword spotting on xylo audio 2 (2024). 2406.15112.
- [55] Shrestha, S. B., Timcheck, J., Frady, P., Campos-Macias, L. & Davies, M. Efficient video and audio processing with loihi 2. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 13481–13485, DOI: 10.1109/ICASSP48485.2024.10448003 (2024).
- [56] Chen, Z. et al. On-off neuromorphic ising machines using fowler-nordheim annealers (2024). 2406.05224.
- [57] Pierro, A. et al. Solving qubo on the loihi 2 neuromorphic processor (2024). 2408.03076.
- [58] Davies, M. et al. Loihi: A neuromorphic manycore processor with on-chip learning. IEEE Micro 38, 82–99, DOI: https://doi.org/10.1109/MM.2018.112130359 (2018).
- [59] Mayr, C., Hoeppner, S. & Furber, S. Spinnaker 2: A 10 million core processor system for brain simulation and machine learning (2019). https://doi.org/10.48550/arXiv.1911.02385.
- [60] Speck. https://www.synsense.ai/products/speck/. Accessed: 2023-04-03.
- [61] Innatera’s Spiking Neural Processor (SNP). www.innatera.com/snp.pdf. Accessed: 2023-11-24.
- [62] Levy, M. Innatera’s Spiking Neural Processor - brain-like architecture targets ultra-low power ai. https://www.innatera.com/innatera-mpr-2021.pdf (2021). Accessed: 2023-12-18.
- [63] TOP500. Top500. https://www.top500.org/ (2023).
- [64] Banbury, C. et al. Mlperf tiny benchmark. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (2021).
- [65] MLCommons. Mlperf inference policies. https://github.com/mlcommons/inference_policies/ (2023).
- [66] Green500. Green500. https://www.top500.org/lists/green500/ (2023).
- [67] MLCommons. Mlcommons power working group. https://mlcommons.org/en/groups/best-practices-power/ (2023).
- [68] Heittola, T., Mesaros, A. & Virtanen, T. Acoustic scene classification in dcase 2020 challenge: generalization across devices and low complexity solutions. In Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020), 56–60 (2020).
- [69] Koch, T., Berthold, T., Pedersen, J. & Vanaret, C. Progress in mathematical programming solvers from 2001 to 2020. EURO Journal on Computational Optimization 10, 100031, DOI: https://doi.org/10.1016/j.ejco.2022.100031 (2022).
- [70] Aimone, J. B. et al. A review of non-cognitive applications for neuromorphic computing. Neuromorphic Computing and Engineering 2, 032003, DOI: https://doi.org/10.1088/2634-4386/ac889c (2022).
- [71] Wurtz, J. et al. Industry applications of neutral-atom quantum computing solving independent set problems (2024). 2205.08500.
- [72] Ke, W., Khoei, M. & Muir, D. Neurobench: Dcase 2020 acoustic scene classification benchmark on xyloaudio 2 (2024). 2410.23776.
- [73] Synsense xylo. https://www.synsense.ai/products/xylo/. Accessed: 2023-04-03.
- [74] D-Wave Systems Inc. D-Wave Samplers software package (2018).
- [75] Rhodes, O. et al. spynnaker: A software package for running pynn simulations on spinnaker. Frontiers in Neuroscience 12, DOI: https://doi.org/10.3389/fnins.2018.00816 (2018).
- [76] SynSense. Samna. https://www.synsense.ai/products/samna/ (2024).
- [77] Pedersen, J. E. et al. Neuromorphic intermediate representation: A unified instruction set for interoperable brain-inspired computing. Nature Communications 15, 8122, DOI: 10.1038/s41467-024-52259-9 (2024).
- [78] Stewart, T. C., DeWolf, T., Kleinhans, A. & Eliasmith, C. Closed-loop neuromorphic benchmarks. Frontiers in Neuroscience 9, DOI: https://doi.org/10.3389/fnins.2015.00464 (2015).
- [79] Jeffares, A., Guo, Q., Stenetorp, P. & Moraitis, T. Spike-inspired rank coding for fast and accurate recurrent neural networks. In International Conference on Learning Representations (2022).
- [80] Prophesee. Event-based vision software - metavision intelligence. https://www.prophesee.ai/metavision-intelligence/ (2023).
- [81] Makin, J. G., O’Doherty, J. E., Cardoso, M. M. B. & Sabes, P. Superior arm-movement decoding from cortex with a new, unsupervised-learning algorithm. J. Neural Eng. 15, DOI: https://doi.org/10.1088/1741-2552/aa9e95 (2018).
- [82] Ansmann, G. Efficiently and easily integrating differential equations with JiTCODE, JiTCDDE, and JiTCSDE. Chaos 28, 043116, DOI: https://doi.org/10.1063/1.5019320 (2018).
- [83] Lin, T.-Y., Goyal, P., Girshick, R., He, K. & Dollár, P. Focal loss for dense object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 318–327, DOI: https://doi.org/10.1109/TPAMI.2018.2858826 (2020).
- [84] Hymel, S. et al. Edge impulse: An mlops platform for tiny machine learning (2023). 2212.03332.
- [85] Muir, D., Bauer, F. & Weidel, P. Rockpool documentaton, DOI: 10.5281/zenodo.3773845 (2019).
## Additional information
### Competing interests statement
The NeuroBench benchmark framework was developed collaboratively and specifically to allow for as objective and applicable comparison as possible. The selection of initial benchmark tasks reflect the authors’ research interests, which includes commercial interests for companies. These do not affect our results in any way, nor the value of our contributions. The benchmark harness and framework are open-source and intended to be further extended by the community over time.