<?xml version="1.0" encoding="UTF-8"?>
<article xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="1.4" article-type="research-article" xml:lang="en"><front><journal-meta><journal-title-group><journal-title xml:lang="ru">Вестник Волгоградского государственного университета. Серия 2. Языкознание</journal-title></journal-title-group><journal-id journal-id-type="issn">1998-9911</journal-id><journal-id journal-id-type="eissn">2409-1979</journal-id></journal-meta><article-meta><article-id pub-id-type="doi">10.15688/jvolsu2.2024.4.9</article-id><title-group><article-title xml:lang="ru">Возможности изучения сочетаемости и устойчивости лексических единиц статистическими методами (на примере глагола take)</article-title><trans-title-group xml:lang="en"><trans-title>Combinability and Stability Analysis of Lexical Units by Statistical Methods (Exemplified by the Verb Take)</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name><surname>Матыцина</surname><given-names>Марина Станиславовна</given-names></name><name-alternatives><name xml:lang="ru"><surname>Матыцина</surname><given-names>Марина Станиславовна</given-names></name><name xml:lang="en"><surname>Matytcina</surname><given-names>Marina</given-names></name></name-alternatives><xref ref-type="aff" rid="aff1"/><email>lipmarina@gmail.com</email><contrib-id contrib-id-type="orcid">0000-0001-6102-4397</contrib-id></contrib><contrib contrib-type="author"><name><surname>Прохорова</surname><given-names>Ольга Николаевна</given-names></name><name-alternatives><name xml:lang="ru"><surname>Прохорова</surname><given-names>Ольга Николаевна</given-names></name><name xml:lang="en"><surname>Prokhorova</surname><given-names>Olga</given-names></name></name-alternatives><xref ref-type="aff" rid="aff2"/><email>prokhorova@bsu.edu.ru</email><contrib-id contrib-id-type="orcid">0000-0001-9441-819X</contrib-id></contrib><contrib contrib-type="author"><name><surname>Чекулай</surname><given-names>Игорь Владимирович</given-names></name><name-alternatives><name xml:lang="ru"><surname>Чекулай</surname><given-names>Игорь Владимирович</given-names></name><name xml:lang="en"><surname>Chekulai</surname><given-names>Igor</given-names></name></name-alternatives><xref ref-type="aff" rid="aff2"/><email>chekulai@bsu.edu.ru</email><contrib-id contrib-id-type="orcid">0000-0001-8599-1699</contrib-id></contrib><aff-alternatives id="aff1"><aff><institution xml:lang="en">Lipetsk State Technical University (Lipetsk, Russian Federation)</institution></aff><aff><institution xml:lang="ru">Липецкий государственный технический университет  (Липецк,, Российская Федерация)</institution></aff></aff-alternatives><aff-alternatives id="aff2"><aff><institution xml:lang="en">Belgorod State National Research University (Belgorod, Russian Federation)</institution></aff><aff><institution xml:lang="ru">Белгородский государственный национальный исследовательский университет (Белгород,, Российская Федерация)</institution></aff></aff-alternatives></contrib-group><pub-date pub-type="epub" iso-8601-date="2024-08-21"><day>21</day><month>08</month><year>2024</year></pub-date><volume>23</volume><issue>4</issue><fpage>106</fpage><lpage>118</lpage><history><date date-type="received" iso-8601-date="2024-03-01"><day>01</day><month>03</month><year>2024</year></date><date date-type="accepted" iso-8601-date="2024-05-13"><day>13</day><month>05</month><year>2024</year></date></history><permissions><license><license-p xml:lang="ru">CC BY 4.0</license-p></license></permissions><abstract xml:lang="ru"><p>Статья посвящена вопросам определения устойчивой сочетаемости слов в речи с применением различных мер ассоциации на примере лингвистического корпуса. Актуальность исследования обусловлена существующей в лингвистике потребностью углубления знаний о факторах, детерминирующих формирование устойчивых отношений элементов внутри словосочетания. В качестве источника избран English Web Corpus (enTenTen) и его подкорпусы. Материалом для анализа послужили биграммы двухсловного сочетания: глагола take с соседним словом. Наряду с критическим рассмотрением мер, используемых для установления связности слов, описан характер отношений между элементами коллокации. Особое внимание уделено сравнению коллокаций в подкорпусах, содержащих тексты разных жанров и тематики. Проанализировано более 100 биграмм, извлеченных посредством мер ассоциации t-score, MI-score и Log Dice. Установлено, что показатели меры t-score различаются в изучаемых подкорпусах, показывают зависимость полученных данных от размера подкорпусов. Делается вывод о том, что вычисление степени устойчивости ассоциативной связи биграмм глагола take, основанное только на этом показателе, невозможно. Данные, полученные с помощью мер MI-score и Log Dice, свидетельствуют о незначительной разнице между подкорпусами, что демонстрирует независимость таких показателей от размера корпуса. Выявлено, что вариативный характер отношений между элементами коллокации заключается в зависимости степени связности слов в словосочетании от частоты их встречаемости в текстах разных жанров, регистров и модальности. М.С. Матыциной подготовлен общий план исследования, осуществлен сбор необходимой информации из корпуса. О.Н. Прохоровой разработана методика анализа, выполнено обобщение материала. И.В. Чекулаем интерпретированы результаты проведенной научной работы.</p></abstract><trans-abstract xml:lang="en"><p>This article is devoted to the issues related to the definition of stable word combinability in speech. The research relevance is sustained by the existing need in profound linguistic knowledge about the factors that determine the formation of stable relationships between the elements of a word combination. The English Web Corpus (enTenTen) and its subcorpora are chosen as the source. The authors consider bigrams of a two-word combination: the verb take with an adjacent word. In addition to a critical examination of the measures used to determine word cohesion, the nature of the relationships between collocation elements is analysed. Particular attention is paid to the comparison of collocations in subcorpora, which contain texts of different genres and topics. More than 100 bigrams obtained through the association measures t-score, MI-score and Log Dice are analysed. The t-score measure differs across the investigated subcorpora, which demonstrates the correlation of the findings with the size of the subcorpora. It is concluded that it is not possible to determine the degree of stability of the associative relationship in the bigrams of the verb take based on this measure alone. The data obtained using the MI-score and Log Dice measures show little difference between subcorpora, demonstrating their independence of the corpus size. The variable nature of the relationships between the collocation elements has been revealed to lie in the dependency of the degree of coherence of words in a word combination on the frequency of their occurrence in the texts of different genres, registers and modalities. Special attention is given to the issue of identifying the degree of effectiveness of the measures in extracting verb collocations and their application to specific professional tasks.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>лингвистический корпус</kwd><kwd>подкорпус</kwd><kwd>коллокация</kwd><kwd>меры ассоциации</kwd><kwd>English Web Corpus (enTenTen)</kwd><kwd>t-score</kwd><kwd>MI-score</kwd><kwd>Log Dice</kwd></kwd-group><kwd-group xml:lang="en"><kwd>linguistic corpus</kwd><kwd>subcorpus</kwd><kwd>collocation</kwd><kwd>measures of association</kwd><kwd>English Web Corpus (enTenTen)</kwd><kwd>t-score</kwd><kwd>MI-score</kwd><kwd>Log Dice</kwd></kwd-group></article-meta></front><back><ref-list><ref id="ref1"><mixed-citation xml:lang="ru">Захаров В. П., Хохлова М. В., 2010. Анализ эффективности статистических методов выявления коллокаций в текстах на русском языке // Компьютерная лингвистика и интеллектуальные технологии : тр. Междунар. конф. «Диалог-2010» (Бекасово, 26–30 мая 2010 г.). М. : Изд-во РГГУ. Вып. 9 (16). С. 137–143.</mixed-citation></ref><ref id="ref2"><mixed-citation xml:lang="ru">Захаров В. П., 2005. Корпусная лингвистика. СПб. : Изд-во СПбГУ. 48 c.</mixed-citation></ref><ref id="ref3"><mixed-citation xml:lang="ru">Мельчук И. А., 1960. О терминах «устойчивость» и «идиоматичность» // Вопросы языкознания. № 4. С. 73–80.</mixed-citation></ref><ref id="ref4"><mixed-citation xml:lang="ru">Филимонов Д. Ю., Светлов А. В., Горбань О. А., Косова М. В., Шептухина Е. М., 2020. Автоматизация процесса метаразметки архивных документов // Математическая физика и компьютерное моделирование. Т. 23, № 4. С. 56–68.</mixed-citation></ref><ref id="ref5"><mixed-citation xml:lang="ru">Шамне Н. Л., Ребрина Л. Н., 2015. Глагольные коллокации памяти в германских СМИ // В мире научных открытий. № 7–8 (67). С. 3097–3108.</mixed-citation></ref><ref id="ref6"><mixed-citation xml:lang="ru">Baker P., 2006. Using Corpora in Discourse Analysis. L. : Bloomsbury Academic. 280 p.</mixed-citation></ref><ref id="ref7"><mixed-citation xml:lang="ru">Burchfield R. W., 1996. The New Fowler’s Modern English Usage. Oxford : Oxford University Press. 864 p.</mixed-citation></ref><ref id="ref8"><mixed-citation xml:lang="ru">Durrant P., Schmitt N., 2009. To What Extent Do Native and Non-Native Writers Make Use of Collocations? // IRAL-International Review of Applied Linguistics in Language Teaching. Vol. 47, iss. 2. P. 157–177. DOI:10.1515/iral.2009.007</mixed-citation></ref><ref id="ref9"><mixed-citation xml:lang="ru">Evert S., 2008. Corpora and Collocations // Corpus Linguistics: An International Handbook. Berlin : Mouton de Gruyter. P. 1212–1248. DOI:10.1515/9783110213881.2.1212</mixed-citation></ref><ref id="ref10"><mixed-citation xml:lang="ru">Henriksen B., 2013. Research on L2 Learners’ Collocational Competence and Development – A Progress Report // L2 Vocabulary Acquisition, Knowledge and Use New Perspectives on Assessment and Corpus Analysis. [S.l.] : [s.n.]. P. 29–56.</mixed-citation></ref><ref id="ref11"><mixed-citation xml:lang="ru">Hill J., 2000. Revising Priorities: From Grammatical Failure to Collocational Success // Teaching Collocation: Further Developments in the Lexical Approach. Hove : LTP. P. 47–67.</mixed-citation></ref><ref id="ref12"><mixed-citation xml:lang="ru">Hunston S., 2002. Corpora in Applied Linguistics. Cambridge : Cambridge University Press. 241 p. DOI:10.1017/CBO9781139524773</mixed-citation></ref><ref id="ref13"><mixed-citation xml:lang="ru">Hunston S., Laviosa S., 2000. Corpus Linguistics. Birmingham : School of English : CELS. 146 p.</mixed-citation></ref><ref id="ref14"><mixed-citation xml:lang="ru">O’Keeffe, A., McCarthy M., Carter R., 2007. From Corpus to Classroom: Language Use and Language Teaching. Cambridge : Cambridge University Press. 332 p.</mixed-citation></ref><ref id="ref15"><mixed-citation xml:lang="ru">McEnery T., Hardie A., 2011. Corpus Linguistics: Method, Theory and Practice. Cambridge : Cambridge University Press. 312 p.</mixed-citation></ref><ref id="ref16"><mixed-citation xml:lang="ru">Sinclair J., 1991. Corpus, Concordance, Collocation. Oxford : Oxford University Press. 179 p.</mixed-citation></ref><ref id="ref17"><mixed-citation xml:lang="ru">Siyanova A., Schmitt N., 2008. L2 Learner Production and Processing of Collocation: A Multi-Study Perspective // Canadian Modern Language Review. Vol. 64, № 3. P. 429–458. DOI: 10.3138/cmlr.64.3.429</mixed-citation></ref><ref id="ref18"><mixed-citation xml:lang="ru">Smadja F., McKeown K. R., Hatzivassiloglou V., 1996. Translating Collocations for Bilingual Lexicons: A Statistical Approach // Computational Linguistics. Vol. 22, № 1. P. 1–38.</mixed-citation></ref><ref id="ref19"><mixed-citation xml:lang="ru">Stubbs M., 1995. Collocations and Semantic Profiles: On the Cause of the Trouble with Quantitative Studies // Functions of Language. Vol. 2, iss. 1. P. 23–55. DOI: 10.1075/fol.2.1.03stu</mixed-citation></ref><ref id="ref20"><mixed-citation xml:lang="ru">Wolter B., Gyllstad H., 2011. Collocational Links in the L2 Mental Lexicon and the Inuence of L1 Intralexical Knowledge // Applied Linguistics. Vol. 32, iss. 4. P. 430–449. DOI: 10.1093/applin/amr011</mixed-citation></ref><ref id="ref21"><mixed-citation xml:lang="ru">EWC – English Web Corpus (enTenTen). URL: https://www.sketchengine.eu/ententen-english-corpus/ (дата обращения: 20.12.2021)</mixed-citation></ref><ref id="ref22"><mixed-citation xml:lang="ru">The Concise Oxford Dictionary of Linguistics. Oxford : Oxford University Press, 2014. 443 p.</mixed-citation></ref><ref id="ref23"><mixed-citation xml:lang="en">Zakharov V.P., Khokhlova M.V., 2010. Analiz effektivnosti statisticheskikh metodov vyyavleniya kollokatsiy v tekstakh na russkom yazyke [Analysis of the Effectiveness of Statistical Methods for Identifying Collocations in Russian Texts]. Kompyuternaya lingvistika i intellektualnye tekhnologii: tr. Mezhdunar. konf. «Dialog-2010» (Bekasovo, 26–30 maya 2010 g.) [Computer Linguistics and Intelligent Technologies. Proceedings of the International Conference “Dialogue-2010” (Bekasovo, May 26–30, 2010)]. Moscow, Izd-vo RGGU, iss. 9 (16), pp.137-143.</mixed-citation></ref><ref id="ref24"><mixed-citation xml:lang="en">Zakharov V.P., 2005. Korpusnaya lingvistika [Corpus Linguistics]. Saint Petersburg, Izd-vo SPbGU. 48 p.</mixed-citation></ref><ref id="ref25"><mixed-citation xml:lang="en">Melchuk I.A., 1960. O terminakh «ustoychivost» i «idiomatichnost» [On Terms “Sustainability” and “Idiomaticity”]. Voprosy yazykoznaniya [Topics in the Study of Language], no. 4, pp. 73-80.</mixed-citation></ref><ref id="ref26"><mixed-citation xml:lang="en">Filimonov D.Yu., Svetlov A.V., Gorban O.A., Kosova M.V., Sheptukhina E.M., 2020. Avtomatizatsiya protsessa metarazmetki arkhivnykh dokumentov [Automation of the Process of Meta-Labelling of Archival Documents]. Matematicheskaya fizika i kompyuternoe modelirovanie [Mathematical Physics and Computer Simulation], vol. 23, no. 4, pp. 56-68.</mixed-citation></ref><ref id="ref27"><mixed-citation xml:lang="en">Schamne N.L., Rebrina L.N., 2015. Glagolnye kollokatsii pamyati v germanskikh SMI [Verb Collocations of Memory in German Media]. V mire nauchnykh otkrytiy [In the World of Scientific Discoveries], no. 7-8 (67), pp. 3097-3108.</mixed-citation></ref><ref id="ref28"><mixed-citation xml:lang="en">Baker P., 2006. Using Corpora in Discourse Analysis. London, Bloomsbury Academic. 280 p.</mixed-citation></ref><ref id="ref29"><mixed-citation xml:lang="en">Burchfield R.W., 1996. The New Fowler’s Modern English Usage. Oxford, Oxford University Press. 864 p.</mixed-citation></ref><ref id="ref30"><mixed-citation xml:lang="en">Durrant P., Schmitt N., 2009. To What Extent Do Native and Non-Native Writers Make Use of Collocations? IRAL-International Review of Applied Linguistics in Language Teaching, vol. 47, iss. 2, pp. 157-177. DOI: 10.1515/iral.2009.007</mixed-citation></ref><ref id="ref31"><mixed-citation xml:lang="en">Evert S., 2008. Corpora and Collocations. Corpus Linguistics: An International Handbook. Berlin, Mouton de Gruyter, pp. 1212-1248. DOI: 10.1515/9783110213881.2.1212</mixed-citation></ref><ref id="ref32"><mixed-citation xml:lang="en">Henriksen B., 2013. Research on L2 Learners’ Collocational Competence and Development – A Progress Report. L2 Vocabulary Acquisition, Knowledge and Use New Perspectives on Assessment and Corpus Analysis. S. l., s. n. P. 29-56.</mixed-citation></ref><ref id="ref33"><mixed-citation xml:lang="en">Hill J., 2000. Revising Priorities: From Grammatical Failure to Collocational Success. Teaching Collocation: Further Developments in the Lexical Approach. Hove, LTP, pp. 47-67.</mixed-citation></ref><ref id="ref34"><mixed-citation xml:lang="en">Hunston S., 2002. Corpora in Applied Linguistics. Cambridge, Cambridge University Press. 241 p. DOI: 10.1017/CBO9781139524773</mixed-citation></ref><ref id="ref35"><mixed-citation xml:lang="en">Hunston S., Laviosa S., 2000. Corpus Linguistics. Birmingham, School of English, CELS. 146 p.</mixed-citation></ref><ref id="ref36"><mixed-citation xml:lang="en">O’Keeffe A., McCarthy M., Carter R. From Corpus to Classroom: Language Use and Language Teaching. Cambridge, Cambridge University Press. 332 p.</mixed-citation></ref><ref id="ref37"><mixed-citation xml:lang="en">McEnery T., Hardie A., 2011. Corpus Linguistics: Method, Theory and Practice. Cambridge, Cambridge University Press. 312 p.</mixed-citation></ref><ref id="ref38"><mixed-citation xml:lang="en">Sinclair J., 1991. Corpus, Concordance, Collocation. Oxford, Oxford University Press. 179 p.</mixed-citation></ref><ref id="ref39"><mixed-citation xml:lang="en">Siyanova A., Schmitt N., 2008. L2 Learner Production and Processing of Collocation: A Multi-Study Perspective. Canadian Modern Language Review, vol. 64, no. 3, pp. 429-458. DOI: 10.3138/cmlr.64.3.429</mixed-citation></ref><ref id="ref40"><mixed-citation xml:lang="en">Smadja F., McKeown K.R., Hatzivassiloglou V., 1996. Translating Collocations for Bilingual Lexicons: A Statistical Approach. Computational Linguistics, vol. 22, no. 1, pp. 1-38.</mixed-citation></ref><ref id="ref41"><mixed-citation xml:lang="en">Stubbs M., 1995. Collocations and Semantic Profiles: On the Cause of the Trouble with Quantitative Studies. Functions of Language, vol. 2, iss. 1, pp. 23-55. DOI: 10.1075/fol.2.1.03stu</mixed-citation></ref><ref id="ref42"><mixed-citation xml:lang="en">Wolter B., Gyllstad H., 2011. Collocational Links in the L2 Mental Lexicon and the Inuence of L1 Intralexical Knowledge. Applied Linguistics, vol. 32, iss. 4, pp. 430-449. DOI: 10.1093/applin/amr011</mixed-citation></ref><ref id="ref43"><mixed-citation xml:lang="en">English Web Corpus (enTenTen). URL: https://www.sketchengine.eu/ententen-english-corpus/ (accessed Dec. 20, 2021)</mixed-citation></ref><ref id="ref44"><mixed-citation xml:lang="en">The Concise Oxford Dictionary of Linguistics. Oxford, Oxford University Press, 2014. 443 p.</mixed-citation></ref></ref-list></back></article>
