This repository contains code and artifacts for our LREC 2026 work on post-hoc Procrustes alignment for Urdu-English cross-lingual retrieval and QA.
We evaluate multilingual embedding models before and after Procrustes alignment on three axes:
- Geometric alignment (cosine distance between parallel sentence pairs)
- Cross-lingual retrieval (Recall@1/3/5)
- Downstream QA quality (RAGAS + EM/F1)
pipeline.ipynb: Main experiment pipeline (data prep, alignment, retrieval)pipeline_2.ipynb: Extended experiments (RAGAS + generation evaluation)Cross_Lingual_Embeddings_LREC_2026.pdf: Paper draftdata/README.md: Dataset access and preparation notesresults/: Structured metric exports for tables/figures
- MiniLM retrieval Recall@1: 0.3871 -> 0.4059 (after alignment)
- LaBSE retrieval Recall@1: 0.3024 -> 0.4273 (after alignment)
- RAGAS faithfulness/context metrics improve across evaluated models
- Generation EM/F1 gains are strongest for weaker pre-alignment models
- Install dependencies:
pip install -r requirements.txt
- Run notebooks in order:
pipeline.ipynbpipeline_2.ipynb
- Export tables to
results/as CSV for paper-ready artifacts.
Squad.csvis intentionally excluded from git due GitHub size limits (>100MB).- Do not commit secrets (
.env, API keys, tokens). - Some RAGAS runs may log
LLMDidNotFinishException; report this in limitations.
See CITATION.cff.