J. Eur. Opt. Society-Rapid Publ. 22, 27( 2026) 281
Fig. 7. Magnetic domain patterns under deterministic initial state, for flip sizes:( a) 1 1,( b) 2 2,( c) 3 3,( d) 4 4,( e) 5 5 and( f) 10 10 lm 2. pffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi N x þ N y
l ¼ tan b
where N x and N y are the numbers of neurons along the horizontal and vertical axes.
Our analysis shows that the input-to-hidden separation d 1, evaluated over a range of values reveals d 1 = 3.0mmas providing optimal classification performance. With d 1 fixed, varying the hidden-to-output spacing indicates that d 2 = 0.5 mm optimizes both accuracy and loss convergence. Further analysis, with d 1 fixed and d 2 set according to the fully connected equation for different flip sizes( Fig. 8), showed that at non-optimal values, accuracy is only weakly affected by interlayer distance, whereas the number of neurons predominantly influences performance. These distances are consistent with diffraction-limited coupling, where firstorder diffracted light efficiently transmits between layers.
3.3 Comparative insights into MCM and SGD-based BP
We evaluated both MCM and SGD with BP in terms of training loss under identical conditions on 5,000 training and 10,000 testing images. BP was trained for 50 epochs with a batch size of 50, totaling 5,000 trials in total, while MCM used the same training and testing sets over 80,000
d n iterations. As shown in Figure 9, MCM starts from higher initial losses, over 50 for deterministic initialization, and around 10 for random initial state. Whereas BP begins near 0.7 and converges faster per trial. Within the first trials, MCM reaches a loss profile nearly identical to BP, indicating that despite its slower start, it ultimately achieves comparable convergence. MCM theoretically requires more iterations( 80,000 trials) and about 20 hours, far exceeding BP, which completes the training in roughly 20 min.
Crucially, MCM permits direct online adaptation within the physical system, an approach that is not accessible to BP owing to its reliance on gradient propagation and discrete activations. Using our optical platform, training was performed on a single-layer MO-D 2 NN. The experimental parameters were aligned with the theoretical computations outlined above, including the inter-layer distances( 3.0 mm, 0.5 mm) and the image size( 100 100 lm 2), ensuring direct correspondence between simulation and implementation. Input images were generated by illuminating a linearly polarized laser beam( k = 532 nm) onto a chromium-on-glass photomask fabricated via photolithography and maskless patterning( MX-1240, Japan Science Engineering Co., Ltd.), where handwritten digit patterns chemically etched. The bismuth-gallium-substituted yttrium iron garnet thin film, served as the hidden-layer owing to its large Faraday effect and high optical transparency.