About

We are a research team founded in 1997

PELCRA (Polish and English Language Corpora for Research and Applications) is a research group at the Institute of English Studies, University of Łódź, Poland, which was originally started in 1997. The group’s main research interests and activities include corpus and computational linguistics, natural language processing, information retrieval, and information extraction. Since 1997, the corpus resources and tools for both English and Polish corpus data have been successfully applied in academic and technological research, as well as in the domains of language teaching, translation, lexicography, and language processing. A list of the corpora, search engines, language processing tools, and electronic dictionaries developed by different members of our group can be found here.

Learn more
Product feature 1
SpokesBiz comprises 8 distinct subcorpora, including a collection of Internet podcasts, biographical interviews, student presentations on academic and popular topics, thematic discussions, spontaneous conversations and more.
Product feature 2
A total of 504 speakers from different regions of Poland, with different educational backgrounds (from primary education to doctoral degrees) and of different ages participated in the formation of the corpus. All these factors contribute to the high diversity and representativeness of the corpus.
Product feature 3
The corpus data were automatically transcribed and time-aligned and subsequently manually corrected. The transcripts are also punctuated using guidelines developed for the DiaBiz corpus.
Product feature 1
The corpus data were collected between May 2021 and April 2022. The phone-call interactions between 5 agents with professional experience in call center settings and 191 participants assuming the role of customers were recorded via the Genesys PureCloud platform.
Product feature 2
DiaBiz can serve as a source of training and evaluation data for a wide range of tasks, such as speech recognition and transcript formatting, speaker diarization, conversational analytics, modelling of dialog systems, and many more.
Product feature 3
The transcripts are also punctuated in order to enhance their clarity. They were first punctuated automatically and subsequently manually corrected by, for which purpose punctuation guidelines that reflect the characteristics of spoken language were developed.

Partners

We cooperate with the best

We welcome all opportunities for cooperation with academic institutions and business partners in various industries alike. Interested? Feel free to contact us.

See all partners