Practical aspects of deploying large language models in local networks
DOI:
https://doi.org/10.34169/2414-0651.2025.1(45).84-96Keywords:
local large language models, confidential information processing, client-server architecture, LLM frameworks, local network, Retrieval-Augmented Generation, data security, server configurationAbstract
The article examines practical aspects of deploying large language models (LLMs) in local networks for processing confidential information. The authors analyze the capabilities of frameworks such as LM Studio, Ollama, Anything LLM and Koboldcpp within different client-server architecture scenarios. Particular attention is paid to functional limitations, technical nuances of configuration, as well as the efficiency and security of utilizing these systems to address tasks in military and scientific domains. Specifically, the article describes a combination of Anything LLM and Ollama frameworks, which ensures integration with knowledge bases through the Retrieval-Augmented Generation (RAG) mechanism, enabling secure data processing in a local environment.
The advantages of architecture based on the Koboldcpp framework, which supports model customization and integration with GPUs to accelerate computations, are detailed. Challenges associated with using LM Studio to create a server infrastructure are also explored, particularly due to its limited API functionality and the complexity of access configuration. The article emphasizes that the optimal choice for locally deploying LLMs in a client-server architecture is the combination of Anything LLM as a user interface and Koboldcpp as a server platform. This system provides flexibility, resource efficiency and high data security, which is particularly critical for applications in the field of national security.
Additionally, the article highlights the importance of testing and configuring the chosen frameworks to ensure system reliability, scalability and compatibility with various user scenarios. The potential use of such architectures for tasks such as real-time intelligence analysis, automation of operational commands and reporting processes in military contexts is also discussed. These solutions are presented as practical and cost-effective approaches to rapidly integrate artificial intelligence technologies into everyday military and scientific activities.
Downloads
References
Slyusar , V. (2024). Local large language models for confidential information processing. Weapons and Military Equipment, 44(4), 79–91.
https://doi.org/10.34169/2414-0651.2024.4(44).79-91
OpenAI. GPT-4 technical report. arXiv, 2024. [Online]. Available: https://arxiv.org/abs/2303.08774.
R. Marino. Fast Analysis of the OpenAI O1-Preview Model in Solving Random K-SAT Problem: Does the LLM Solve the Problem Itself or Call an External SAT Solver. arXiv preprint arXiv:2409.11232, 2024. [Online]. Available: https://arxiv.org/abs/2409.11232.
LM Studio. 2024. [Online]. Available: https://lmstudio.ai/.
Ollama. 2024. [Online]. Available: https://ollama.com/.
AnythingLLM. The all-in-one AI application, 2024. [Online]. Available: https://anythingllm.com/.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems. Vol. 33. Pp. 9459—9474. [Online]. Available: https://arxiv.org/abs/2005.11401.
A. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. Singh Chaplot, D. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. Renard Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, W. E. Sayed. Mistral 7B. 2023. 9 p. https://arxiv.org/pdf/2310.06825.pdf.
Parthasarathy, V.B., Zafar, A., Khan, A. & Shahid, A. (2024). The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities. arXiv preprint arXiv:2408.13296. [Online]. Available: https://arxiv.org/abs/2408.13296.
Slyusar Vadym. Large language models (LLM) in the military area. XXIII Intern. Scientific and Technical Conf. «Artificial Intelligence and Intellegent Systems» (AIIS’2023). 10 − 11 October, 2023. https://doi.org/10.13140/RG.2.2.30196.94086.
Slyusar Vadym. Reducing the Cognitive Burden of a Soldier with the Help of Personal AI and LLM Assistant. The LCGDSS Human System Integration (HSI) symposium, 12 January 2024. https://doi.org/10.13140/RG.2.2.10264.57605/1.
Slyusar, V.I., Kondratenko, Y.P., Shevchenko, A.I. & Yeroshenko, T.V. (2024). Some Aspects of Artificial Intelligence Development Strategy for Mobile Technologies. J. of Mobile Multimedia. Vol. 20_3. Pp. 525—554.
https://doi.org/10.13052/jmm1550-4646.2031.
Vakulenko, M. & Slyusar, V. (2024) Automatic Smart Subword Segmentation for the Reverse Ukrainian Physical Dictionary Task. Proc. of the Modern Data Science Technologies Workshop (MoDaST-2024). Lviv. Ukraine. May 31 – June 1. Pp. 59—73.
Slyusar, V., Sotnyk, V., & Chepkov, I. (2024). WEAPON SYSTEMS RESEARCH METHODOLOGY: TEORETICAL AND PRACTICAL ASPECTS. Weapons and Military Equipment, 43(3), 3–8.
https://doi.org/10.34169/2414-0651.2024.3(43).3-8
Koboldcpp. 2024. [Online]. Available at: https://github.com/LostRuins/koboldcpp.
TroyDoesAI/Agent-Flow-Phone_Demo_3GB_RAM. 2024. [Online]. Available at: https://huggingface.co/TroyDoesAI/Agent-Flow-Phone_Demo_3GB_RAM.
YorkieDev. LM Studio Server Examples. GitHub. Jun. 2024. [Online]. Available at: https://github.com/YorkieDev/lmstudioservercodeexamples. [Accessed: Sep. 28, 2024].
Issa2k23. LMStudio integration. GitHub. Jun. 18. 2024. [Online]. Available at: https://github.com/open-webui/open-webui/discussions/3280. [Accessed: Sep. 28, 2024].
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Вадим Слюсар,Іван Гусаковський

This work is licensed under a Creative Commons Attribution 4.0 International License.