278
J. Eur. Opt. Society-Rapid Publ. 22, 27( 2026)
were used for most implementations. All experiments are repeated with multiple random seeds to insure statistical robustness.
2.3 Flip-size and initialization strategies in MCM optimization
Fig. 2. Structural overview of the MO-D 2 NN model featuring a single hidden layer and a polarizer.
rescaled to 100 100 pixels. Input images were propagated through the simulated optical network, and the output plane was segmented into ten-class specific detection areas corresponding to digits 0 9, each measuring 11 11 lm 2. The predicted class was determined based on output plane intensity, providing a direct mapping from output pixels to their expected digit.
2.2.3 Loss function and evaluation
Convergence was tracked by loss reduction and test accuracy, with hyperparameters( flip size, iterations) set from preliminary runs. Loss was evaluated as the cross-entropy between target labels and normalized output intensities within each class regions, and classification accuracy was computed on the test set. Binary weights were updated via MCM proposals, accepting flips if loss did not increase or remained unchanged.
2.2.4
Training protocol
Training proceeded with fixed flip-size MCM updates, described above, for up to 80,000 iterations. At each step, a subset of binary weights awas proposed for flipping and losswasevaluatedforbothconfigurations. For each configuration we report convergence curves( loss vs iterations), and test accuracy vs. iterations. Optimization was entirely gradient-free, driven by stochastic binary search. Backpropagation baseline was included under identical data partitions for reference. When evaluating the loss, the algorithm computes DL, ensuring a consistent loss reduction without requiring explicit gradient computation.
2.2.5 Implementation details
Numerical simulations were conducted on a standard workstation-AI equipped with NVIDIA RTX-A6000( 48 GB of graphical processing unit-GPU, a Core i9-10980XE-18 cores / 36 thread of central processing unit-CPU, and 128 GB( 32GB4) of random-access memory-RAM), running on Ubuntu 20.04 LTS operating system. Python( v3.10.9), NumPy( 1.23.2), and TensorFlow( v2.12.0) framework
To further investigate the impact of the initialization states and the update strategy on the optimization process, we evaluated multiple flip-sizes of the network weights. Initialization plays a key role in guiding the network toward optimal configurations, while the update strategy governs its exploration during each iteration. The choice of flip size shapes this process: smaller flips promote finer adjustments and more stable convergence, whereas larger flips enable broader exploration but may induce instability.
In this study, we examined two initialization configuration schemes:
- Deterministic initial state: where all the magnetic domains are initially aligned at the same direction(+ 1). Then number of weight flips were applied per iteration.
- Random initial state: where domains of the network are initialized stochastically at different directions(+ 1 or �1), following the same flip size sequence.
By systematically varying flip sizes( F s ¼ f1 1; 2 2; 3 3; 4 4; 5 5; 10 10 lm 2 g) within the two initialization states, we assessed their impact on convergence behavior, stability, accuracy, and overall performance. This analysis reveals the influence of the initial magnetic configuration on network performance and provides a comprehensive understanding on the interplay between initialization, update strategy, and flip size on convergence and optimization efficiency, offering valuable insights for selecting parameters that ensure robust network optimization.
3 Results and discussion
The performance of the MCM optimization was evaluated by comparing different fixed flip size strategies under different initialization configurations.
3.1 Effect of initialization and flip size on convergence
Figure 3 shows the evolution of training loss and accuracy convergence for deterministic and random initializations with a flip-size of 1 1 lm 2 over the same number of iterations( 80,000). In both cases, the loss decreased steadily without divergence, achieving monotonic or plateau-like loss reduction, consistent with the non-worsening acceptance rule that rejects unfavorable updates and allows equal-loss moves. However, their convergence characteristics differ underscoring the strong influence of initialization on MCM optimization.
In the case of deterministic initialization, where all weights set to + 1, loss exhibits a slower initial decrease in