276
J. Eur. Opt. Society-Rapid Publ. 22, 27( 2026)
widely employed within Nasa research, supporting Bayesian inference and state estimation with neural fields [ 7, 8 ], uncertainty quantification( UQ) in aerodynamic simulations [ 9 ], and recent Bayesian / UQ-initiatives aimed at assessing spacecraft performance under uncertainty and improving predictive modeling [ 10, 11 ]. By employing stochastic sampling with probabilistic acceptance, MCMs explore high-dimensional search spaces without relying on explicit derivatives. This makes them intrinsically compatible with the discrete, uncertain, and partially observable nature of physical neural systems. Their probabilistic structure also mitigates local minima, supports asynchronous updates, and minimizes unnecessary weight rewrites, key advantages for devices with limited endurance, like MO domain-switching networks. These features make MCM particularly well-suited for in-situ training, offering online adaptability and reconfigurability during operation, unlike conventional BP, which is generally limited to offline optimization and static inference.
This paper presents a theoretical framework of a Monte Carlo-based algorithm for online learning in MO-D 2 NN. The convergence behavior, classification performance on image recognition tasks( Handwritten digits, MNIST), and the underlying optimization strategy are examined through a generalized analysis. The proposed approach and its practical feasibility are confirmed were experimentally validated for online learning in MO-D 2 NN, as reported in our prior study [ 12 ]. The results provide new insights into Monte Carlo-driven physical neural networks and highlight its potential as a scalable scheme for next-generation neural computing. Although slower than conventional gradientbased methods in offline training, the approach is well suited to real-time physical implementation, where backpropagation cannot be directly applied. However, our research has primarily focused on MO-D 2 NN, where the proposed MCM-based learning has been validated, yet it is not restricted to MO-D 2 NN and can be extended to general diffractive deep neural networks( D 2 NNs). Because training relies exclusively on forward propagation combined with MCM updates, it is gradient-free and enables direct implementation in systems with discrete, or non-differentiable elements.
2 Monte Carlo optimization algorithmic approach
2.1 General MCM framework
Metropolis-inspired MCM optimization algorithm employed here is a gradient-free search strategy, in which binary weights of the diffractive network are iteratively updated through stochastic proposals and loss-guided acceptance. Each trainable configuration is represented by a binary weight vector v e { �1, 1 } N, whereN denotes the number of trainable domains, as diffractive elements. The optimization of these binary weights is formulated as a stochastic search process directly guided by the system loss function. At each iteration, a new configuration v ' is generated by flipping a predefined number of weight elements.
At each iteration, the optimization proceeds as follows:
- Initialization: The algorithm begins by randomly initializing a binary weight vector v( 0) 2 { 0, 1 } N, representing the initial magnetic domain configuration.
- Sampling generation: A new state v 0 is generated by flipping a subset of entries in the current magnetic domain in v, corresponding to local modifications of the domain configuration.
- Loss evaluation: The corresponding loss L( v 0) is computed by propagating the input through MO-D 2 NN and comparing the predicted distribution with the true label distribution. The change in loss in then obtained as:
L ¼ Lðv 0 Þ�LðvÞ
- Acceptance rule: The update rule follows a deterministic Metropolis acceptance criterion. The candidate state is deterministically accepted if DL < 0, and the magnetic domain pattern is updated accordingly. Otherwise, the algorithm retains the previous configuration.
If L < 0 update v ðtþ1Þ ¼ v 0; else;
retain v ðtþ1Þ ¼ v ðÞ t
- Iteration and Termination: The steps of sampling, evaluation, and acceptance are repeated until convergence, or a predefined maximum number of iterations( MCiterations) is reached, hereafter referred to as N max.
The efficiency of MCM optimization depends on its ability to explore the weight configurations within the hidden layer, which grow exponentially with network size. At each iteration, the network processes input data, computes the loss relative to the target, proposes random domain updates in the magnetic-optical layer, and applies weight flips to generate candidate configurations aimed at reducing the loss. Candidate configurations are accepted or rejected according to a Metropolis criterion, enabling derivative-free optimization( see the conceptual flowchart, Fig. 1).
2.2 Application of MCM in practice
The MCM optimization was applied to train a fully connected MO-D 2 NN with binary phase encoded weights for image classification tasks. The network operates as a sequence of optical transformations, where each diffractive layer modulates the incident light. The resulting intensity distribution is measured at the output plane.
The forward light propagation was calculated numerically, accounting for the magneto-optical( MO) effect by separately propagating the right and left circularly polarized light. The MO effect, measured as polarization rotation and ellipticity, arises from differences in refractive index and extinction coefficient for the two circular polarizations. Free-space propagation was computed using the bandlimited angular spectrum method [ 13 ], derived from the