ICON 2026 · Shared Task A COIL-D Project Initiative

Translate Indian languages directly into Indian languages.

COILD-INDIC-MT is a shared task on direct, non-pivoted machine translation between neighbouring Indian languages — no English or Hindi leg anywhere in the pipeline. 12 translation directions, three language families, seven writing systems.

Bidirectional pairs
  • Assamese ↔ Bodo
  • Kannada ↔ Malayalam
  • Bengali ↔ Odia
  • Gujarati ↔ Marathi
  • Dogri ↔ Punjabi
  • Urdu ↔ Sindhi
Indo-Aryan · Dravidian · Tibeto-Burman
Motivation

Pivoting through English loses what makes these languages distinct.

Machine translation between Indian languages is, in practice, almost never between Indian languages. Where an Indic–Indic direction is supported at all, it is typically realised by pivoting through English or Hindi — lossy in exactly the places neighbouring Indian languages are rich: honorific agreement, classifier constructions, kinship terminology, and case and converb morphology. It also discards genuine areal proximity, and makes translation for a Bodo or Dogri speaker contingent on the state of English resources.

Why COIL-D

The COIL-D (Centre of Indian Language Data) project builds a unified repository of Indian language resources, sets benchmarking standards, and advances Machine Translation and NLP technologies for Human Language Technology applications. COILD-INDIC-MT is COIL-D's shared task focused exclusively on direct, non-pivoted translation between regional Indian languages.

A deliberately varied set

Rather than maximising language count, each pair isolates a different relationship: a cross-family pair convergent through contact (Assamese–Bodo); within-family pairs at differing degrees of relatedness (Bengali–Odia, Gujarati–Marathi, Kannada–Malayalam); a pair with an extremely low-resource source (Dogri–Punjabi); and a pair sharing a script family but diverging lexically (Urdu–Sindhi).

Shared task

12 directions, twelve languages

All twelve languages are listed in the Eighth Schedule of the Constitution of India. Participants may enter any subset of the six bidirectional language pairs.

# Bidirectional pair Family (src ↔ tgt) Script (src ↔ tgt)
1Assamese ↔ BodoIndo-Aryan ↔ Tibeto-BurmanBengali–Assamese ↔ Devanagari
2Kannada ↔ MalayalamDravidian ↔ DravidianKannada ↔ Malayalam
3Bengali ↔ OdiaIndo-Aryan ↔ Indo-AryanBengali ↔ Odia
4Gujarati ↔ MarathiIndo-Aryan ↔ Indo-AryanGujarati ↔ Devanagari
5Dogri ↔ PunjabiIndo-Aryan ↔ Indo-AryanDevanagari ↔ Gurmukhi
6Urdu ↔ SindhiIndo-Aryan ↔ Indo-AryanPerso-Arabic ↔ Perso-Arabic

Objectives

  • Facilitate effective communication across India's diverse linguistic communities through direct machine translation, without relying on a pivot language.
  • Encourage multilingual and transfer learning across the community of Indian languages.
  • Provide a curated, high-quality multilingual parallel corpus to the community.
  • Help neighbouring communities communicate directly using machine translation.
  • Advance systems that preserve the linguistic and cultural diversity of close communities.
  • Establish a common evaluation platform for Indic-centric multilingual MT.
The corpus

~30,000 parallel sentences per direction

COILD-INDIC-MT provides a carefully curated, high-quality parallel corpus covering all six participating language pairs. Training and development data are released to registered participants; the blind test set (source-side only) is released separately during the evaluation window and remains hidden until then.

Corpus splits — per language pair

Training set28,000
Development set1,000
Practice test set1,000
Total per pair~30,000

Across all six language pairs

Training set168,000
Development set6,000
Practice test set6,000
Total corpus~180,000
Access policy. The complete dataset (training, development, and practice test sets) is released only to registered participants. Access to the official blind test set is also restricted to registered teams, who must complete registration and agree to the dataset usage policy before obtaining access.
Requesting the dataset. Once registered, the training, development, and practice test sets are distributed via a Hugging Face dataset request. Request access on Hugging Face — coming soon → Use the same email address you registered with — only the team lead may submit the request, on behalf of the whole team.
Participation

Two submission tracks

Each participating team must submit at least one system to be eligible for submitting a system description paper. Teams may enter any subset of the six language pairs.

Track 1

Primary (Constrained)

Systems must be trained using only the datasets released by the organizers.

Up to 3 runs per language pair, with one run designated as the primary submission.

Track 2

Open (Unconstrained)

Systems may utilize additional publicly available datasets or external resources.

Up to 3 runs per language pair, with one run designated as the primary submission.

All submissions must comply with the shared task guidelines and data usage policy specified by the organizers.

Timeline

Key dates

The gold marker tracks today's position on the line automatically, based on your device's date. All deadlines are 23:59 AoE (Anywhere on Earth) unless noted otherwise.

22 Aug 2026

Shared task webpage goes live & registrations start

5 Sept 2026

Training set released (registered participants only)

10 Oct 2026

Evaluation (test) set released (registration closes)

13 Oct 2026

Final date for run submissions

20 Oct 2026

Task results and rankings published

05 Nov 2026

Participant system papers due

15 Nov 2026

Acceptance decisions communicated

25 Nov 2026

Camera-ready working notes

Get involved

Register your team

Registration is open as of 22 August 2026. Registered teams get access to the training and development data, and later the blind test set.

The registration link will be added here as soon as it's live — check back, or watch this page.

1

Register your team

Registration form link will appear on this page.
2

Agree to the data usage policy

Required before training/development data access is granted.
3

Get training data (25 Aug 2026)

Development and practice test sets released alongside it.
4

Submit runs by 13 Oct 2026

Up to 3 runs per direction, per track — one marked primary.
Organizing committee

Organizers

Kshetrimayum Boynao Singh

Kshetrimayum Boynao Singh

IIT Patna, India
Deepak Kumar

Deepak Kumar

IIT Patna, India
Avinash Kumar

Avinash Kumar

IIT Patna, India
Nitin Kumar Mishra

Nitin Kumar Mishra

IIT Delhi, India
Palash Pratim Dutta

Palash Pratim Dutta

IIT Guwahati, India
Ashwini Vaidya

Ashwini Vaidya

IIT Delhi, India
Sanasam Ranbir Singh

Sanasam Ranbir Singh

IIT Guwahati, India
Samit Bhattacharya

Samit Bhattacharya

IIT Guwahati, India
Shad Akhtar

Shad Akhtar

IIIT Delhi, India
Poonam Bansal

Poonam Bansal

IGDTUW, Delhi, India
Muralikrishna SN

Muralikrishna SN

MIT, Manipal, India
Tanmoy Chakraborty

Tanmoy Chakraborty

IIT Delhi, India
Asif Ekbal

Asif Ekbal

IIT Patna, India
Questions about the task?