Beyond GPUs: Can Optical and Neuromorphic Chips Replace Backprop?
GPUs are hitting efficiency limits for AI. Here's why researchers are exploring optical computing, neuromorphic chips, and gradient-free training instead.

Why are people questioning GPUs at all?
GPUs won the AI hardware race because they were good at the thing deep learning needed: matrix multiplication at scale. But the argument now circulating among hardware researchers is that GPUs are optimized for a narrow, historically contingent path, convolutional networks, then transformers, then backpropagation, then memory-hungry attention, and that path is running into diminishing returns. Compute efficiency measured in gigaflops per joule has plateaued over roughly the last two years even as models keep growing. The industry responded by throwing more memory bandwidth and capacity at the problem instead of more raw efficiency, because attention scales quadratically with sequence length and that demand had to go somewhere. The result is chips exquisitely tuned for one architecture and one training algorithm, which makes them fast today and potentially a dead end tomorrow.
TL;DR
- GPU efficiency has stalled: gigaflops-per-joule gains have flattened over the past couple of years, even as model sizes and memory demands keep climbing.
- Nvidia’s hardware evolved around transformers, not raw compute efficiency, shifting focus toward memory capacity and bandwidth once GPT-era models showed how memory-hungry attention mechanisms are.
- The human brain runs on about 20 watts and performs feats current AI hardware can’t match per joule, which is why neuromorphic and brain-inspired approaches keep resurfacing as a research direction.
- Backpropagation doesn’t map onto biological neurons, since the brain is largely feedforward and doesn’t appear to run the backward error-passing step that backprop requires.
- Optical computing promises near-free computation in the form of light, since beams of light can perform certain operations with very low energy, but light is hard to store, making memory and training in optics a real challenge.
- Zero-order (gradient-free) optimization methods, like SPSA, can train models without computing gradients at all, trading some efficiency at scale for architectures that don’t need backward passes.
- None of these alternatives beat GPUs today, but each attacks a different bottleneck (energy, memory, or the training algorithm itself), and combining them is where the more speculative research is heading.
Built like a system. Not vibe-coded.
Remy manages the project — every layer architected, not stitched together at the last second.
What’s actually wrong with the GPU plus backprop combination?
The case against the current stack isn’t that GPUs are slow. It’s that the whole system, chip architecture, training algorithm, and model design, co-evolved together in a way that locks in specific tradeoffs. Convolutional networks dominated roughly from 2012 (the AlexNet era) through around 2019-2020, and GPU design during that window favored raw floating-point throughput because CNNs weren’t particularly memory or bandwidth bound. When transformer-based language models arrived and demonstrated that attention mechanisms demand far more memory capacity and bandwidth, GPU makers pivoted hardware generations toward stacking more memory rather than squeezing more flops per watt. That’s a reasonable response to the models people wanted to train, but it means the hardware roadmap is now a product of what transformers need, not necessarily what intelligence needs.
Backpropagation itself has a biological problem. It requires a model to run a forward pass, then propagate error signals backward through the exact same connections to update weights. Real neurons don’t appear to do this. The brain is largely feedforward, and for a neuron to “run backward” through a synapse it would need some mechanism to reverse its own signaling, which isn’t how biological firing works. Attempts to make Hebbian-style learning (where connections strengthen based on correlated activity) substitute for backprop run into what’s called the weight transport problem: the backward pass would need to use weights that exactly mirror the forward pass’s weights, which has no clear biological mechanism. Some workarounds exist if neurons on both sides observe the same signals, but what results is closer to a cyclic graph of local updates than anything resembling true backpropagation.
How does optical computing attack this problem?
Data centers already move data as light. Fiber optic connections handle nearly all long-distance and much short-distance data transfer because optics is extremely energy-efficient for communication. The inefficiency creeps in at the boundaries: light gets converted to electrical signals so a chip can compute on it, then converted back to light to move to the next stage, then back to electrons again, and so on. Each conversion between optical and electronic domains burns energy, and the actual computation, sitting in electronic transistors, is where the overwhelming majority of power gets consumed.
The optical computing pitch is to skip the conversions and do the computation itself in light. Certain mathematical operations, including types of matrix multiplication relevant to neural networks, can be performed by light propagating through engineered optical structures, essentially using physics to do the arithmetic instead of switching transistors. Energy cost for that kind of computation can be extremely low since photons don’t dissipate heat the way electrical current through resistive circuits does. Diffusion-based image generation, where light propagation is programmed to perform the generative steps, has been explored as a proof of concept for this kind of approach.
The catch is storage. Light doesn’t sit still. Holding a bit of information in an optical medium for later use is far harder and more expensive than storing a bit in electronic memory. That means optical systems are promising for raw computation but struggle with anything that needs persistent state, which is a problem for standard neural network training that depends on storing and updating weights repeatedly.
What is neuromorphic computing, and why does it matter?
Neuromorphic computing tries to build chips that mimic the brain’s structure directly in analog circuitry rather than simulating neurons in software on digital hardware. The appeal is energy efficiency: biological brains perform complex computation on roughly the power of a household light bulb, a figure that remains far out of reach for digital AI systems running comparable workloads.
Part of what neuromorphic designs try to capture is structural, not just computational. The brain isn’t just a dense web of connections. It contains repeated structures called cortical columns, essentially localized assemblies of neurons, surrounded by heavy inhibitory signaling that lets each assembly learn somewhat independently of its neighbors. That combination of local processing and strong inhibition is a structural feature that standard neural networks, and the GPUs built to run them, don’t really replicate. Neuromorphic chips attempt to build that kind of localized, massively parallel, low-power architecture in silicon or other analog substrates.
Can you train neural networks without backpropagation?
Yes, though with tradeoffs. Zero-order optimization methods update a model using only forward passes, no gradients computed through backward differentiation at all. One such method, called SPSA (simultaneous perturbation stochastic approximation), works by perturbing a model’s parameters in a random direction, running two forward passes (one with a small positive perturbation, one with a small negative one), and using the difference between the resulting losses to decide how much to update in that direction. Repeat that process enough times with enough random perturbations and the model converges, without ever computing a gradient in the traditional sense.
This approach has real advantages: it works on loss landscapes that aren’t smooth or differentiable, where gradient-based methods can get stuck in local minima, and it doesn’t require storing activations for a backward pass. The tradeoff is that the noise in the gradient estimate gets worse as model size grows, scaling unfavorably with parameter count. That makes zero-order methods impractical for training very large models directly, since the number of perturbations needed to get a usable signal becomes enormous.
One proposed workaround is to shard a large model into many small experts, cluster training data, and train each small model independently with zero-order methods before combining them with a lightweight router at inference time, similar in spirit to mixture-of-experts but built from many independently trained small models rather than one jointly trained network. Because each shard is small, the gradient estimate noise stays manageable without needing as many perturbations. This kind of approach only makes sense, though, if the cost of each forward pass drops dramatically, which is exactly the opening that optical or neuromorphic hardware could provide.
Is any of this ready to replace GPUs?
One coffee. One working app.
You bring the idea. Remy manages the project.
Not yet. Optical computing systems face real engineering limits around storing information and handling the non-differentiable, noisy behavior of physical light systems. Neuromorphic chips are still largely experimental and haven’t matched the raw throughput of digital accelerators for mainstream workloads. Zero-order optimization methods don’t scale efficiently to today’s largest models using today’s hardware. Each approach solves one piece of the puzzle, energy cost, biological plausibility, or training algorithm, without yet solving all three together. What makes the research area interesting is that these pieces could in principle combine: cheap optical or neuromorphic forward passes paired with gradient-free training that doesn’t need a backward pass at all, removing the core assumption GPUs were built around in the first place.
Frequently Asked Questions
What’s the main criticism of GPUs for AI workloads?
The criticism isn’t raw speed, it’s that gigaflop-per-joule efficiency has plateaued recently while hardware design has shifted toward adding memory capacity and bandwidth to serve transformer attention mechanisms rather than improving fundamental compute efficiency.
Why doesn’t the brain use backpropagation?
Backpropagation requires a backward pass through the same connections used in the forward pass, which would require neurons to somehow fire in reverse. Biological neurons appear to operate in a largely feedforward manner, and proposed biological alternatives like Hebbian learning run into the weight transport problem, where backward and forward weights would need to match exactly with no clear mechanism to enforce that.
What is optical computing in simple terms?
It means performing mathematical computation, like matrix multiplication, using light propagating through engineered structures rather than electrical signals moving through transistors. It can be extremely energy-efficient for computation but struggles with storing information for later use.
What is zero-order optimization?
It’s a family of training methods, including one called SPSA, that update model parameters using only forward passes and the resulting change in loss, without ever computing a gradient through backpropagation. It handles rough, non-differentiable loss landscapes well but becomes noisier and less efficient as model size increases.
Are optical or neuromorphic chips available commercially today?
These remain largely research-stage technologies. The techniques discussed, optical matrix computation and brain-inspired analog circuits, have been demonstrated in academic and early-stage research settings rather than deployed at the scale of commercial GPU training clusters.