ISBN-13: 9781495992872 / Angielski / Miękka / 2014 / 58 str.
ISBN-13: 9781495992872 / Angielski / Miękka / 2014 / 58 str.
The purpose of this document is to describe the best practices that personnel from the National Institute of Standards and Technology (NIST) have developed and implemented to efficiently and effectively capture two-way, free-form speech-to-speech audio dialogues within recording studios. These dialogues, produced to support the development and evaluation of machine translation technologies, are conducted by English and foreign language speakers conversing with one another in their native languages through the mediation of an interpreter. NIST personnel have collected over 500 h of bilingual audio data sets encompassing more than 1100 dialogues across three unique language pairs (English/Iraqi-Arabic, English/Dari, and English/Pashto) since it became involved in this work in 2007.