Pivoting through English loses what makes these languages distinct.
Machine translation between Indian languages is, in practice, almost never between Indian languages. Where an Indic–Indic direction is supported at all, it is typically realised by pivoting through English or Hindi — lossy in exactly the places neighbouring Indian languages are rich: honorific agreement, classifier constructions, kinship terminology, and case and converb morphology. It also discards genuine areal proximity, and makes translation for a Bodo or Dogri speaker contingent on the state of English resources.
Why COIL-D
The COIL-D (Centre of Indian Language Data) project builds a unified repository of Indian language resources, sets benchmarking standards, and advances Machine Translation and NLP technologies for Human Language Technology applications. COILD-INDIC-MT is COIL-D's shared task focused exclusively on direct, non-pivoted translation between regional Indian languages.
A deliberately varied set
Rather than maximising language count, each pair isolates a different relationship: a cross-family pair convergent through contact (Assamese–Bodo); within-family pairs at differing degrees of relatedness (Bengali–Odia, Gujarati–Marathi, Kannada–Malayalam); a pair with an extremely low-resource source (Dogri–Punjabi); and a pair sharing a script family but diverging lexically (Urdu–Sindhi).
12 directions, twelve languages
All twelve languages are listed in the Eighth Schedule of the Constitution of India. Participants may enter any subset of the six bidirectional language pairs.
| # | Bidirectional pair | Family (src ↔ tgt) | Script (src ↔ tgt) |
|---|---|---|---|
| 1 | Assamese ↔ Bodo | Indo-Aryan ↔ Tibeto-Burman | Bengali–Assamese ↔ Devanagari |
| 2 | Kannada ↔ Malayalam | Dravidian ↔ Dravidian | Kannada ↔ Malayalam |
| 3 | Bengali ↔ Odia | Indo-Aryan ↔ Indo-Aryan | Bengali ↔ Odia |
| 4 | Gujarati ↔ Marathi | Indo-Aryan ↔ Indo-Aryan | Gujarati ↔ Devanagari |
| 5 | Dogri ↔ Punjabi | Indo-Aryan ↔ Indo-Aryan | Devanagari ↔ Gurmukhi |
| 6 | Urdu ↔ Sindhi | Indo-Aryan ↔ Indo-Aryan | Perso-Arabic ↔ Perso-Arabic |
Objectives
- Facilitate effective communication across India's diverse linguistic communities through direct machine translation, without relying on a pivot language.
- Encourage multilingual and transfer learning across the community of Indian languages.
- Provide a curated, high-quality multilingual parallel corpus to the community.
- Help neighbouring communities communicate directly using machine translation.
- Advance systems that preserve the linguistic and cultural diversity of close communities.
- Establish a common evaluation platform for Indic-centric multilingual MT.
~30,000 parallel sentences per direction
COILD-INDIC-MT provides a carefully curated, high-quality parallel corpus covering all six participating language pairs. Training and development data are released to registered participants; the blind test set (source-side only) is released separately during the evaluation window and remains hidden until then.
Corpus splits — per language pair
Across all six language pairs
Two submission tracks
Each participating team must submit at least one system to be eligible for submitting a system description paper. Teams may enter any subset of the six language pairs.
Primary (Constrained)
Systems must be trained using only the datasets released by the organizers.
Up to 3 runs per language pair, with one run designated as the primary submission.
Open (Unconstrained)
Systems may utilize additional publicly available datasets or external resources.
Up to 3 runs per language pair, with one run designated as the primary submission.
All submissions must comply with the shared task guidelines and data usage policy specified by the organizers.
Key dates
The gold marker tracks today's position on the line automatically, based on your device's date. All deadlines are 23:59 AoE (Anywhere on Earth) unless noted otherwise.
Shared task webpage goes live & registrations start
Training set released (registered participants only)
Evaluation (test) set released (registration closes)
Final date for run submissions
Task results and rankings published
Participant system papers due
Acceptance decisions communicated
Camera-ready working notes
Register your team
Registration is open as of 22 August 2026. Registered teams get access to the training and development data, and later the blind test set.
The registration link will be added here as soon as it's live — check back, or watch this page.
Register your team
Registration form link will appear on this page.Agree to the data usage policy
Required before training/development data access is granted.Get training data (25 Aug 2026)
Development and practice test sets released alongside it.Submit runs by 13 Oct 2026
Up to 3 runs per direction, per track — one marked primary.