Publications
A paradigmatic perspective on semantic transparency in derivation
Semantic transparency in derivation – the extent to which the meaning of a derived word can be inferred from its morphological structure and constituents – is typically assessed through the relationship between a derivative and its base. Recent research has emphasized the importance of analyzing derived words within broader networks of morphologically related lexemes, raising the question of how transparency should be conceptualized when derivational paradigms are considered. To address this issue, we distinguish two dimensions of transparency corresponding to orthogonal paradigmatic relationships and operationalize them distributionally. Process typicality (T_proc) captures the semantic consistency between a derivative and the other words formed through the same derivational process, whereas family relatedness (T_fam) captures its consistency with the other members of its derivational family. Applying both measures to over 32,000 French derivatives from Démonette-2.0, we show that the two dimensions are empirically dissociable and shaped by distinct linguistic factors. Process typicality is primarily influenced by properties of derivatives and processes (frequency, productivity, lexical ambiguity), whereas family relatedness is primarily determined by properties of derivational families (size, age, ambiguity of family members) – with several predictors showing opposite effects across the two dimensions. These findings support a paradigmatic, multidimensional account of semantic transparency.
Beyond free variation: a probabilistic hierarchy in cumulative intensifying prefixation
Cumulative intensifying prefixation – in which multiple synonymous prefixes stack onto a single base, as in It. super-iper-bello ‘super-hyper-nice’ – has traditionally been treated as a domain of free variation, on the grounds that the prefixes are interchangeable. Drawing on 312 multi-prefix adjectival types from five corpora of Italian, we show that interchangeability in the grammar does not entail free variation in usage. For each of the fifteen pairs formed by the six prefixes examined (arci-, extra-, iper-, stra-, super-, ultra-), one order is preferred, and the fifteen preferences fit together without contradiction into a single ranking – a configuration unrelated preferences would produce only 2.2% of the time. Resampling shows the ranking to be gradient rather than rigid, its extremes stable and its interior uncertain. We then ask what it tracks, comparing two information-theoretic measures: base entropy, a proxy for productivity, and paradigmatic surprisal, a proxy for expressive markedness. Prefixes with higher productivity reliably appear farther from the base, while expressive markedness shows no such effect. A morphological paradigm can therefore acquire systematic internal organization through usage alone, without any grammatical constraint enforcing it.
Systematic or arbitrary? Quantifying distributional structure in affix rivalry through Italian intensifying prefixes
A central question in research on morphological competition is to what extent the distribution of rival affixes is governed by systematic distributional regularities or by arbitrary lexeme-specific conventions – and how these two forces can be disentangled empirically. This study addresses this question through the case of six Italian intensifying prefixes – arci-, extra-, iper-, stra-, super-, and ultra- – in adjectival derivation, a domain of evaluative morphology where rivalry has received almost no systematic attention. Using a dataset of 2,700 adjectival derivative tokens annotated for formal (phonological, syntactic, semantic) and usage-based (base–prefix association, base age, frequency, polarity) properties, we compare three machine-learning models that represent increasingly abstract views of what conditions prefix choice: (i) a model that includes item-specific base-prefix association scores, capturing entrenched pairwise conventions; (ii) a model that relies exclusively on generalizable base-level properties; and (iii) a type-based model that removes token-frequency effects entirely. The first two models perform comparably, indicating that coarse-grained base properties recover nearly all of the distributional signal captured by item-specific associations. The type-based model, by contrast, shows a sharp performance drop with a restructured predictor hierarchy, suggesting that much of the systematicity in intensifying prefix rivalry is usage-driven rather than purely categorical. These findings point to a two-tier architecture of affix rivalry in which systematic usage-based constraints structure the primary distributional signal, while item-specific entrenchment plays a secondary, complementary role.
How rivalrous are rivals? A distributional semantic approach to gradient competition in evaluative morphology
Affix rivalry is a form of morphological competition in which multiple affixes stand in a many-to-one relationship with a shared or closely related function. Despite growing interest, evaluative morphology has remained largely peripheral to this line of research. Through a distributional semantic study of six Italian intensifying prefixes (arci-, extra-, iper-, stra-, super-, and ultra-), we argue that rivalry is gradient rather than binary: morphological processes compete to a greater or lesser degree depending on how much their semantic transformations overlap. We model both derivational outputs and the semantic transformations they induce as vectors in semantic space, and propose an information-theoretic operationalization of rivalry as the proportion of uncertainty about process identity that remains after observing the semantic transformation. Converging analyses show that the six prefixes are neither semantically equivalent nor categorically distinct: they occupy substantially overlapping regions of semantic space while retaining weak but detectable prefix-specific signatures. Iper-, super-, and ultra- form a particularly tight core, inducing highly similar semantic shifts, and only about 10% of the uncertainty about prefix identity is recoverable from the semantic transformations. The analyses further show that classification accuracy and information- theoretic measures can diverge, underscoring the value of the latter for quantifying rivalry. Overall, rivalry among evaluative prefixes appears to share the gradient organization observed in other derivational domains, differing in degree rather than kind: competing within a narrow functional domain, intensifying prefixes sit toward the high-overlap end of the rivalry continuum, where rivals remain only marginally distinguishable.
Competition in evaluation: An empirical study of rivalry in Italian intensifying prefixation
Morphology is often seen as a domain of rules, yet it is also a domain of choices. This study investigates the phenomenon of rivalry in Italian evaluative morphology, focusing on intensifying adjectival constructions formed with the prefixes arci-, extra-, iper-, stra-, super-, and ultra- (e.g., arcicontento ‘overjoyed’, stracotto ‘overcooked’, superveloce ‘superfast’). While affix rivalry has been extensively studied in non-evaluative derivation, evaluative morphology - particularly the function of intensification - has received limited attention. The present research aims to illuminate how factors such as productivity, semantic overlap, and usage-based variables interact to shape prefix rivalry dynamics and speakers’ choices among competing forms. To this end, given prefix polyfunctionality, a manually annotated large-scale dataset of intensified derivatives formed with the six prefixes was compiled, comprising 48,069 tokens distributed over 3,686 types. Since gradience inherent in rivalry requires nuanced analytical approaches, a comprehensive series of quantitative corpus analyses, distributional semantics techniques, and machine-learning modeling was applied to the dataset to derive statistical generalizations. Initial analyses reveal that iper-, super-, and ultra- are highly productive and share overlapping distributional niches, forming a core cluster of rival prefixes. In contrast, arci- and stra- are largely confined to lexicalized forms, whereas extra- occupies an intermediate, semi-autonomous position. Distributional semantics analyses corroborate these findings, showing substantial overlap among the prefixes, with rivalry strongest among iper-, super-, and ultra-, which tend to induce highly similar semantic shifts and display no significant differences in semantic transparency. Innovative token-based modeling of (extra)linguistic predictors further indicates that base-prefix association, accounting for entrenched pairings, is the strongest predictor of prefix selection, followed by base frequency, age, and polarity, while interpretable machine learning methods shed light on prefix-specific patterns. Unlike prototypical derivational systems, where formal constraints typically prevail, the distribution of Italian intensifying prefixes appears to be strongly shaped by usage-based factors. Collectively, the results reveal a complex, usage-driven system of gradient rivalry among Italian intensifying prefixes. More broadly, they highlight the need for theoretical models that can accommodate the probabilistic dynamics underlying morphological competition. Embracing this non-categorical landscape of coining opportunities represents a promising avenue for future research on rivalry in word-formation.
Deriving semantic classes of Italian adjectives via word embeddings: a large-scale investigation
This paper investigates the application of word embeddings to derive semantic classes for Italian adjectives. Adjectives were clustered using UMAP for dimensionality reduction and K-means for clustering. Semantic categories such as “Relational”, “Descriptive”, “Evaluative”, “Membership”, and “Physical/HealthRelated” were tested by employing predefined prototypical adjectives for each class. The precision and recall of the classification were analyzed, revealing high accuracy for some classes (e.g., “Evaluative”), but challenges in distinguishing more nuanced categories such as “Descriptive”. Furthermore, cluster overlaps were visualized using KDE and quantified using KNN, , highlighting semantic intermingling between groups, especially between the “Descriptive” and “Evaluative” categories. Finally, a comparison with Wordnet’s adjective categories was provided.
InTens – a dataset of Italian intensified derivatives. Description and application in a productivity study
The paper introduces InTens, a dataset of Italian intensified adjectival derivatives formed with six evaluative prefixes, namely arci-, extra-, iper-, stra-, super-, and ultra-. Initially, we delineate the process of data extraction and filtration. Subsequently, we address the polyfunctional characteristics of the evaluative prefixes, with semantic annotation of derivatives into two macro- categories: intensification and non-intensification. After discussing the annotation results, the application of InTens is demonstrated through an investigation of the morphological productivity of the prefixes. The analysis underscores the variability in productivity contingent upon the semantic function of the prefixes, an aspect most often overlooked in productivity research.
Approximation by morphological means: exploring prefixoids kvazi(-), nadri(-), nazovi(-), and pseudo(-) in Croatian.
This study explores the phenomenon of affix rivalry within the domain of morphological approximation in Croatian, focusing on the prefixoids kvazi(-), nadri(-), nazovi(-), and pseudo(-) as they attach to nominal bases. These prefixoids can be classified as privative, as the derivatives they produce do not fully embody the core characteristics conveyed by their morphological bases. To analyze the rivalry among the prefixoids, the study evaluates their productivity, collocational behavior, and distribution across various textual genres, utilizing data from the CLASSLA-web.hr corpus. The findings suggest significant disparities in the productivity and collocational behavior of the prefixoids, with nazovi(-) and kvazi(-) exhibiting the highest productivity and highly overlapping collocational behavior, whereas pseudo(-) and nadri(-) reveal more specialized usage patterns. Additionally, a random sample of 500 tokens per prefixoid is annotated for semantic values. Again, nazovi(-) and kvazi(-) demonstrate substantial overlap, particularly in their mutual application as means for subjective depreciative evaluation, underscoring the insufficiency or pretentiousness of the subject. Nadri(-) is more narrowly focused on legal domains, while pseudo(-), with its proclivity for scientific contexts, remains distinct but conceptually adjacent to kvazi(-) in contexts where imitation is highlighted without necessarily invoking deceit. Overall, the prefixoids present a complex network of interrelationships, yet each prefixoid also establishes a specific niche, balancing between shared semantic roles and distinct, context-dependent uses.
An insight into the Croatian degree modifier paradigm and its clustering profiles
Degree modifiers represent linguistic items employed to alter other elements in relation to their degree. Despite being a well–studied category in English linguistics, degree modifiers in Croatian have received limited attention. This study aims to address this gap by examining a set of Croatian degree modifiers as a part of <degree modifier + adjective> construction. Initially, a corpus analysis is used, and the 29 most frequent degree modifiers of adjectives in the hrWaC corpus are identified. To analyse the examined modifier, we turn to the distributional hypothesis and examine collocational contexts in which modifiers occur. By employing a simple collexeme analysis, we quantify the degree of attraction between a given degree modifier and adjective for each <degree modifier + adjective> construction and its 1000 most frequent adjectival collocates. The results of simple collexeme analysis then serve as input for hierarchical agglomerative cluster analysis, shedding light on the clustering patterns of Croatian degree modifiers based on their favoured collexemes. Simple collexeme analysis reveals itself as successful in filtering out collexems that consistently appear irrespective of the context, proving its superiority over methods relying solely on raw frequencies. The subsequent cluster analysis exposes some discrepancies between the modifiers’ function and their cluster profiling, resulting in clusters lacking functional homogeneity. Nonetheless, certain subclusters demonstrate perfect or almost perfect stability and empirical support, affirming the (near–)synonymy among involved modifiers.
A corpus-based study of maximizer–adjective patterns in Croatian
Maximizers represent a subclass of degree modifiers that convey the highest degree to which a property can be carried out. This paper studies five Croatian near-synonymous maximizers (all meaning “completely, totally”), viz. posve, potpuno, sasvim, skroz, and totalno, as a part of <maximizer + adjective> construction. It is assumed here that analysed pairings act as (semi)-prefabricated units with maximizers that impose particular modes of construal. To analyse the subtle semantic differences of examined maximizers, we shall turn to the distributional hypothesis and examine contexts in which maximizers occur. Using a combination of analytical statistics (collostructional analysis) and multifactorial methods (hierarchical agglomerative cluster analysis and correspondence analysis), we aim to examine similarities (proximities) and differences (distances) between analysed constructions in order to understand intricate relationships among maximizers, fostering valuable insights into their semantics. The findings of this study provide insight into the interplay of the Croatian maximizers and adjectives.
Intensificazione degli aggettivi ai tempi della pandemia: analisi di un corpus di articoli giornalistici
La lingua, il riflesso del mondo colpito dalla pandemia del Covid-19, negli ultimi due anni è stata costretta a mettere in pratica diversi procedimenti linguistici per trasmettere più efficacemente la severità del “nuovo normale”. Un fenomeno linguistico investente non solo il livello morfologico, ma altresì quello semantico-pragmatico, rafforzando oppure indebolendo la forza referenziale di un enunciato, è l’intensificazione. L’intensificazione rappresenta l’insieme eterogeneo delle strategie linguistiche finalizzate a modulare il contenuto proposizionale di un elemento lessicale variandolo per intensità. Benché l’intensificazione si presenti come un fenomeno transcategoriale che può essere associato a tutte le classi grammaticali – ammesso che esse possano essere in qualche modo graduate – sono giustappunto gli aggettivi, parte del discorso intrinsecamente scalare, la categoria lessicale più coinvolta nel suddetto processo. Il presente lavoro propone un approfondimento delle strategie di intensificazione degli aggettivi riguardante il livello della morfologia (prefissazione, suffissazione ed elativo) e della sintassi (modificazione mediante avverbi) su un corpus di 288 articoli giornalistici trattanti i temi della pandemia. Lo studio, oltre alle premesse teoriche, ovvero alla presentazione dei corpora in esame, prevede un’analisi contrastiva approfondita del fenomeno di intensificazione mediante l’aggiunta di avverbi che verrà osservato in relazione alla frequenza dei medesimi avverbi nei corpora itTenTen16 e COMPARE-IT.
Psovke u drami ‘Predstava Hamleta u selu Mrduša Donja’: kontrastivna analiza izvornika i prijevoda na istromletački dijalekt
Psovački izrazi, emotivno nabijeni formulaički stereotipni jezični izrazi kojima se označuju tabuizirani objekti i zbivanja iz izvanjezične zbilje, neizbježna su pojava u svakodnevnoj komunikaciji. Unatoč njihovoj sveprisutnosti, ovaj se jezični fenomen nije često proučavao u hrvatskim jezikoslovnim istraživanjima, što zbog percepcije kako je riječ o manje vrijednome dijelu jezičnih iskaza, što zbog problematike terminološkoga određenja pojma psovke. Nadalje, psovački se izrazi temelje na kulturološkim aspektima društva koje se njima služi pa je njihovo prevođenje jedan od većih izazova traduktološke prakse. Jezik Brešanove drame “Predstava Hamleta u selu Mrduša Donja” (1965) pokazuje visok stupanj inovativnosti, a dijalozi, odnosno sama tematska okosnica djela, travestiran, stilski degradiran i u Dalmatinsku zagoru ubiciran Shakespeareov kanonski tekst, obiluju psovačkim izrazima. U ovome radu primijenit će se kontrastivni pristup i analizirati psovke u spomenutoj grotesknoj tragediji i u njezinu prijevodu na istromletački dijalekt. Psovke iz obiju drama klasificirat će se ovisno o njihovoj strukturi (morfološko-sintaktička razina) i ovisno o domenama ljudskoga života na koje se referiraju (funkcionalno-semantička razina) kako bi se u nastavku taj izdvojeni korpus mogao kontrastivno analizirati. Cilj je rada opisati strategije prijevoda psovki s hrvatskoga jezika na istromletački dijalekt kako bi se razumjeli razlozi odabira određene prijevodne varijante. Također, provedena analiza trebala bi pokazati u kojoj mjeri psovke u prijevodnoj inačici zadržavaju svoj izvorni ilokucijski, odnosno perlokucijski učinak.
