Multi-Layer AI Architecture for Real-Time Detection of Workflow and Data Integrity Failures in Distributed Retail Platforms

Sai Sruthi Puchakayala

Abstract


Distributed retail systems run on heterogeneous data systems whose joint state represents business invariants that cannot be verified in the life-cycle of any one of them․ Silent inconsistencies arise when‚ downstream‚ purchase orders go missing‚ quantities do not match‚ suppliers are blocked‚ or data drifts across systems․ Rule-based monitoring systems and single-system observability do not detect these failures by definition․ Text-to-SQL and retrieval-augmented generation technologies answer these questions without executing the multi-step investigation an incident demands․ We present a two-layer AI architecture that separates deterministic data access and non-deterministic reasoning․ The semantic data access layer‚ implemented as an MCP server‚ defines eight domain tools as well as five templates with natural-language descriptions․ The stateful reasoning layer parses incidents‚ reasons about them‚ selects tools‚ takes investigative steps and derives remediation following a seven-step workflow captured as a knowledge-based process in a Camunda BPMN process engine․ From production data of a large international retailer‚ the mean time to resolution was reduced from 4-8 hours to 1-2 hours․ The average across the 5 classes of anomaly is 77 percent reduced․ Tool description debt is a new class of technical debt that is specific to LLM-based observability systems․ A four-element taxonomy is proposed․


Full Text:

PDF

References


P. Notaro, J. Cardoso, and M. Gerndt, "A survey of AIOps methods for failure management," ACM Trans. Intell. Syst. Technol., vol. 12, no. 6, art. 81, 2021, doi: 10.1145/3483424.

X. Hou, Y. Zhao, S. Wang, and H. Wang, "Model context protocol (MCP): landscape, security threats, and future research directions," arXiv:2503.23278, 2025, doi: 10.48550/arXiv.2503.23278.

M. M. Hasan, H. Li, E. Fallahzadeh, G. K. Rajbahadur, B. Adams, and A. E. Hassan, "Model context protocol (MCP) at first glance: studying the security and maintainability of MCP servers," arXiv:2506.13538, 2025, doi: 10.48550/arXiv.2506.13538.

A. S. Rao and M. P. Georgeff, "BDI agents: from theory to practice," in Proc. 1st Int. Conf. Multi-Agent Systems (ICMAS), 1995, pp. 312–319.

J. Ortiz, V. Torres, and P. Valderas, "Microservice compositions based on the choreography of BPMN fragments: facing evolution issues," Computing, vol. 105, no. 1, pp. 375–416, 2023, doi: 10.1007/s00607-022-01128-8.

W. Cunningham, "The WyCash portfolio management system," ACM SIGPLAN OOPS Messenger, vol. 4, no. 2, pp. 29–30, 1992, doi: 10.1145/157710.157715.

P. Kruchten, R. L. Nord, and I. Ozkaya, "Technical debt: from metaphor to theory and practice," IEEE Software, vol. 29, no. 6, pp. 18–21, 2012, doi: 10.1109/MS.2012.167.

G. Pang, C. Shen, L. Cao, and A. van den Hengel, "Deep learning for anomaly detection: a review," ACM Comput. Surv., vol. 54, no. 2, art. 38, 2022, doi: 10.1145/3439950.

B. Li, X. Peng, Q. Xiang, H. Wang, T. Xie, J. Sun, and X. Liu, "Enjoy your observability: an industrial survey of microservice tracing and analysis," Empirical Software Engineering, vol. 27, no. 1, art. 25, 2022, doi: 10.1007/s10664-021-10063-9.

J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou, "Chain-of-thought prompting elicits reasoning in large language models," in Proc. Adv. Neural Inf. Process. Syst. 35, 2022, pp. 24824–24837, doi: 10.48550/arXiv.2201.11903.

S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, "ReAct: synergizing reasoning and acting in language models," in Proc. 11th Int. Conf. Learn. Representations, 2023, doi: 10.48550/arXiv.2210.03629.

T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, "Toolformer: language models can teach themselves to use tools," in Proc. Adv. Neural Inf. Process. Syst. 36, 2023, doi: 10.48550/arXiv.2302.04761.

L. Wang, C. Ma, X. Feng, et al., "A survey on large language model based autonomous agents," Frontiers of Computer Science, vol. 18, no. 6, 186345, 2024, doi: 10.1007/s11704-024-40231-1.

T. Guo, X. Chen, Y. Wang, et al., "Large language model based multi-agents: a survey of progress and challenges," in Proc. 33rd Int. Joint Conf. Artif. Intell. (IJCAI), 2024, pp. 8048–8057, doi: 10.24963/ijcai.2024/890.

Q. Zhang, C. Fang, Y. Xie, et al., "A survey on large language models for software engineering," arXiv:2312.15223, 2024, doi: 10.48550/arXiv.2312.15223.

Y. Chen, H. Xie, M. Ma, et al., "Automatic root cause analysis via large language models for cloud incidents," in Proc. 19th European Conf. Comput. Syst. (EuroSys), 2024, pp. 674–688, doi: 10.1145/3627703.3629553.

Z. Hong, Z. Yuan, Q. Zhang, H. Chen, J. Dong, F. Huang, and X. Huang, "Next-generation database interfaces: a survey of LLM-based text-to-SQL," arXiv:2406.08426, 2024, doi: 10.48550/arXiv.2406.08426.

X. Liu, S. Shen, B. Li, et al., "A survey of text-to-SQL in the era of LLMs: where are we, and where are we going?" arXiv:2408.05109, 2024, doi: 10.48550/arXiv.2408.05109.

Y. Gao, Y. Xiong, X. Gao, et al., "Retrieval-augmented generation for large language models: a survey," arXiv:2312.10997, 2024, doi: 10.48550/arXiv.2312.10997.

L. Huang, W. Yu, W. Ma, et al., "A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions," ACM Trans. Inf. Syst., 2024, doi: 10.1145/3703155.

P. Sahoo, P. Meharia, A. Ghosh, S. Saha, V. Jain, and A. Chadha, "A comprehensive survey of hallucination in large language, image, video and audio foundation models," in Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 11709–11724, doi: 10.18653/v1/2024.findings-emnlp.685.

D. L. Parnas, "On the criteria to be used in decomposing systems into modules," Communications of the ACM, vol. 15, no. 12, pp. 1053–1058, 1972, doi: 10.1145/361598.361623.

A. Cockburn, "Hexagonal architecture (ports and adapters)," 2005. [Online]. Available: https://alistair.cockburn.us/hexagonal-architecture/

R. Levin, E. Cohen, W. Corwin, F. Pollack, and W. Wulf, "Policy/mechanism separation in Hydra," in Proc. 5th ACM Symp. Operating Systems Principles (SOSP), 1975, pp. 132–140, doi: 10.1145/800213.806531.

M. Attaran, S. Attaran, and B. G. Celik, "The impact of digital twins on the evolution of intelligent manufacturing and Industry 4.0," Advances in Computational Intelligence, vol. 3, no. 11, 2023, doi: 10.1007/s43674-023-00058-y.

A. Samuels, "Examining the integration of artificial intelligence in supply chain management from Industry 4.0 to 6.0: a systematic literature review," Frontiers in Artificial Intelligence, vol. 7, art. 1477044, 2025, doi: 10.3389/frai.2024.1477044.

P. Avgeriou, P. Kruchten, I. Ozkaya, and C. Seaman, "Managing technical debt in software engineering (Dagstuhl seminar 16162)," Dagstuhl Reports, vol. 6, no. 4, pp. 110–138, 2016, doi: 10.4230/DagRep.6.4.110.


Refbacks

  • There are currently no refbacks.


Abava  Кибербезопасность Monetec 2026 СНЭ

ISSN: 2307-8162