Optimizations in metric learning model training
Researchers

Paulo R. Lisboa de Almeida
UFPR

André Grégio
UFPR

Roberto Tomchak
UFPR

Heloísa Viotto
UFPR

Cauê Samonek
UFPR

Thamiris Yamate Fischer
UFPR
Maintaining Difficulty: A Margin Scheduler for Triplet Loss in Siamese Networks Training
This study reveals that, during the training of Siamese Networks using Triplet Margin Ranking Loss, the effective margin is often higher than the hyperparameter’s preset value, which can limit learning. The paper then suggests adapting this value during training; however, unlike previous approaches that focus on specific scenarios, the authors propose adjusting the margin based on the ratio of easy triplets.
In experiments conducted on the LFW, CelebA, CUB-200-2011, and Stanford Cars datasets, the effectiveness of a fixed margin of m = 0.3 was compared against two methods: the proposed Linear Scheduler, which starts with an initial margin of m0 = 0.0 and a growth rate per epoch of sl = 0.01; and the Difficulty Adaptive Margin Scheduler (DAMS), which also starts with an initial margin of m0 = 0.0, but applies a growth rate of sa = 0.01 only in epochs where the ratio of easy triplets exceeds a threshold of t = 0.95. These hyperparameters were selected based on exhaustive testing across multiple configurations.
The results indicate that DAMS is the most suitable method in most scenarios, as its margin growth can plateau at a certain point during training, unlike the linear method, which can be overly aggressive, or the fixed margin, which fails to adapt to the problem.
This work, titled “Maintaining Difficulty: A Margin Scheduler for Triplet Loss in Siamese Network Training” was presented at IJCNN 2026.
How Far is Too Far? Defining the Distance Threshold for Verification Siamese Networks.
In Siamese Networks, two objects are considered to belong to the same class if the distance between them falls below a predefined threshold. This paper proposes determining the threshold (t) for each problem by assuming that the distribution of distances can be approximated by a bimodal function. In this scenario, t is the local minimum between the two modes: positive pairs concentrated at smaller distances and negative pairs at larger ones. Unlike other methods, such as selecting this hyperparameter based on ROC curve analysis, the proposed approach does not require pre-labeled data and is robust against distribution shifts between training, validation, and test sets. Furthermore, the technique performs well on imbalanced data.
The paper presents two algorithms: BimodalVal, which estimates the threshold using only the validation dataset, and BimodalUpd, where t is initialized based on the validation data but can adapt every 64 new distances observed during testing. In experiments conducted on the MNIST, CIFAR10, PKLot, and LFW datasets, these two algorithms were compared against a fixed threshold baseline and the Equal Error Rate (EER) method. The results demonstrate that both proposed algorithms and EER achieve similar accuracy, significantly outperforming the fixed threshold approach.
This work, titled “How Far is Too Far? Defining the Distance Threshold for Verification Siamese Networks”, was presented at SMC 2026.
