Adaptation of machine translation for multilingual information retrieval in the medical domain.

Pavel Pecina, Ondřej Dušek, Lorraine Goeuriot, Jan Hajič, Jaroslava Hlaváčová, Gareth J F Jones, Liadh Kelly, Johannes Leveling, David Mareček, Michal Novák, Martin Popel, Rudolf Rosa, Aleš Tamchyna, Zdeňka Urešová

Artificial Intelligence in Medicine 2014 July

OBJECTIVE: We investigate machine translation (MT) of user search queries in the context of cross-lingual information retrieval (IR) in the medical domain. The main focus is on techniques to adapt MT to increase translation quality; however, we also explore MT adaptation to improve effectiveness of cross-lingual IR.

METHODS AND DATA: Our MT system is Moses, a state-of-the-art phrase-based statistical machine translation system. The IR system is based on the BM25 retrieval model implemented in the Lucene search engine. The MT techniques employed in this work include in-domain training and tuning, intelligent training data selection, optimization of phrase table configuration, compound splitting, and exploiting synonyms as translation variants. The IR methods include morphological normalization and using multiple translation variants for query expansion. The experiments are performed and thoroughly evaluated on three language pairs: Czech-English, German-English, and French-English. MT quality is evaluated on data sets created within the Khresmoi project and IR effectiveness is tested on the CLEF eHealth 2013 data sets.

RESULTS: The search query translation results achieved in our experiments are outstanding - our systems outperform not only our strong baselines, but also Google Translate and Microsoft Bing Translator in direct comparison carried out on all the language pairs. The baseline BLEU scores increased from 26.59 to 41.45 for Czech-English, from 23.03 to 40.82 for German-English, and from 32.67 to 40.82 for French-English. This is a 55% improvement on average. In terms of the IR performance on this particular test collection, a significant improvement over the baseline is achieved only for French-English. For Czech-English and German-English, the increased MT quality does not lead to better IR results.

CONCLUSIONS: Most of the MT techniques employed in our experiments improve MT of medical search queries. Especially the intelligent training data selection proves to be very successful for domain adaptation of MT. Certain improvements are also obtained from German compound splitting on the source language side. Translation quality, however, does not appear to correlate with the IR performance - better translation does not necessarily yield better retrieval. We discuss in detail the contribution of the individual techniques and state-of-the-art features and provide future research directions.

Full text links

We have located links that may give you full text access.

Show additional links to paperHide additional links to paper

PubMed

Add to Saved Papers

Get 1-tap access

Related Resources

Challenges in Septic Shock: From New Hemodynamics to Blood Purification Therapies.Fernando Ramasco et al.Journal of Personalized Medicine 2024 Februrary 4

Molecular Targets of Novel Therapeutics for Diabetic Kidney Disease: A New Era of Nephroprotection.Alessio Mazzieri et al.International Journal of Molecular Sciences 2024 April 4

Perioperative echocardiographic strain analysis: what anesthesiologists should know.Adrian Costescu et al.Canadian Journal of Anaesthesia 2024 April 11

The 'Ten Commandments' for the 2023 European Society of Cardiology guidelines for the management of endocarditis.Michael A Borger, Victoria DelgadoEuropean Heart Journal 2024 April 18

Prevention and treatment of ischaemic and haemorrhagic stroke in people with diabetes mellitus: a focus on glucose control and comorbidities.Simona Sacco et al.Diabetologia 2024 April 17

Essential thrombocythaemia: A contemporary approach with new drugs on the horizon.Francisca Ferrer-Marín et al.British Journal of Haematology 2024 April 9

For the best experience, use the Read mobile app

Get seemless 1-tap access through your institution/university

For the best experience, use the Read mobile app

All material on this website is protected by copyright, Copyright © 1994-2024 by WebMD LLC.
This website also contains material copyrighted by 3rd parties.

By using this service, you agree to our terms of use and privacy policy.

Your Privacy Choices

You can now claim free CME credits for this literature searchClaim now

Get seemless 1-tap access through your institution/university

For the best experience, use the Read mobile app

Adaptation of machine translation for multilingual information retrieval in the medical domain.

Full text links

Related Resources

Trending Papers

For the best experience, use the Read mobile app