WG-5
AI/ML in Analysis
Training and Evaluation Framework for Army Embedding Models
MAJ Jonathan Kasprisin, MAJ David Allen
Transformation Decision Analysis Center
An Army-focused framework develops datasets, fine-tunes embedding models, and evaluates retrieval quality, cost, and latency.
Abstract
Artificial intelligence (AI) workflows commonly use retrieval-augmented generation (RAG) to reduce hallucinations and improve factual accuracy. RAG systems use embedding models to convert text to semantic vectors that can be searched or compared with a user's query. Embedding models are a critical component of RAG-based systems because they influence retrieval quality, computational cost, and latency.
The Transformation Decision Analysis Center (TDAC) currently uses an open source embedding model for a RAG-based system that supports Army analysts. Open source and general-purpose embedding models are free and transparent, but they may struggle to represent domain-specific semantic patterns. Research has demonstrated that general-purpose models can improve performance on domain-specific tasks through fine-tuning. This research develops a framework to train and evaluate embedding models fine-tuned for the Army-domain.
To achieve this, we expand upon existing datasets to develop Army-focused datasets for evaluating embedding models. We then use the evaluation sets to identify performance gaps when using general-purpose embedding models. Based on the observed gaps, we generate a domain-informed synthetic training dataset and fine-tune a candidate embedding model for Army-specific applications. Lastly, the trained model is evaluated against general-purpose baselines to quantify improvements in retrieval effectiveness, cost, and latency.
This research provides TDAC and the Army Operations Research Symposium (AORS) with Army-specific embedding benchmarks, evidence-based guidance on performance–cost trade offs, and a reusable framework for data generation, evaluation, and embedding model tuning.
The Transformation Decision Analysis Center (TDAC) currently uses an open source embedding model for a RAG-based system that supports Army analysts. Open source and general-purpose embedding models are free and transparent, but they may struggle to represent domain-specific semantic patterns. Research has demonstrated that general-purpose models can improve performance on domain-specific tasks through fine-tuning. This research develops a framework to train and evaluate embedding models fine-tuned for the Army-domain.
To achieve this, we expand upon existing datasets to develop Army-focused datasets for evaluating embedding models. We then use the evaluation sets to identify performance gaps when using general-purpose embedding models. Based on the observed gaps, we generate a domain-informed synthetic training dataset and fine-tune a candidate embedding model for Army-specific applications. Lastly, the trained model is evaluated against general-purpose baselines to quantify improvements in retrieval effectiveness, cost, and latency.
This research provides TDAC and the Army Operations Research Symposium (AORS) with Army-specific embedding benchmarks, evidence-based guidance on performance–cost trade offs, and a reusable framework for data generation, evaluation, and embedding model tuning.
Presenters
- MAJ Jonathan KasprisinTransformation Decision Analysis Center
- MAJ David AllenTransformation Decision Analysis Center