RISK RANKING OF AUTHENTICATION EVENTS BASED ON TEMPORAL AND BEHAVIORAL CONTEXT

Authors: Makatov Ye.K., Bishwajeet Pandey, Ismukanova A.N., Glock E.S., Esmagambetova G.K.
IRSTI 81.93.29

Abstract. Not all authentication log events require the same analysis priority; therefore, ranking them using temporal and behavioral context is relevant to information security. This study examines temporal and behavioral features for ranking authentication events by relative risk. The objective is to assess the contribution of behavioral representation to ranking quality and robustness, statistical validity, explainability, and calibration, and to test the technical feasibility of behavioral prioritization on real-world logs. Methods included baseline model comparison, representation ablation, chronological train/validation/test splitting, empirical random-ranking comparison, robustness and statistical analyses, explainability, calibration, low-and-slow temporal-horizon sensitivity, and behavioral feature-group ablation; NDCG@100 was the primary ranking metric.In the synthetic experiment, R2_behavioral_core achieved a higher mean NDCG@100 than R0_minimal (0.126952 vs. 0.071314) and outperformed it in 10 of 12 scenario–model combinations. Across 60 paired observations, mean ΔNDCG@100 was +0.046803 with a 95% confidence interval of [0.035444, 0.057751]. Empirical random-ranking analysis provided an additional reference for interpreting the absolute ranking scores: for Logistic Regression with R2_behavioral_core, NDCG@100 was 0.189052 versus a random mean of 0.010045 in the baseline scenario and 0.145514 versus 0.004943 in the rare-attacks scenario. The behavioral advantage remained condition-dependent and weakened under low-and-slow activity. Explainability and calibration were used to characterize score contributions, explanation stability, calibration quality, and ranking preservation.
For 1,760,511 real-world log events, prioritization scores and unique ranks were calculated from five past-only behavioral features, satisfying all 9/9 temporal/representation integrity requirements. Predictive effectiveness was not evaluated because verified security labels were unavailable. The synthetic results support the incremental ranking value of temporal and behavioral context, while the real-world analysis demonstrates the technical feasibility of implementing the corresponding past-only behavioral prioritization procedure. Confirming predictive effectiveness in operational settings requires temporal external validation using logs with verified security labels and semantic field documentation.

Keywords: authentication events, behavioral analysis, risk ranking, temporal validation, security event prioritization, machine learning, explainability.