Polysemy Disambiguation via Contrastive Semantic Mapping
This research introduces Contrastive Semantic Mapping, a novel framework for Word Sense Disambiguation (WSD), the core NLP task of resolving polysemy—words with multiple context-dependent meanings—to improve downstream applications like machine translation, information retrieval, and sentiment analysis. Unlike traditional knowledge-based or statistical WSD methods that struggle with fine-grained semantic nuances and limited annotated training data, this approach leverages contrastive learning to explicitly model and separate semantic boundaries for different word senses in a high-dimensional embedding space. The framework follows a structured workflow: it starts with preprocessing text corpora to identify polysemous words and their contexts, constructs contrastive positive (same-sense) and negative (different-sense) pairs with stratified sampling and hard-negative mining to reduce bias, uses transformer-based encoders to generate contextual embeddings, and optimizes a contrastive loss function to pull same-sense representations closer while pushing different-sense representations apart. Extensive benchmark testing on standard WSD datasets confirms that this framework outperforms traditional and cutting-edge baseline methods including SVM and standard BERT, with statistically significant improvements in accuracy and F1-score, particularly for fine-grained sense distinctions. It also supports modular updates for emerging new word senses without full model retraining, making it scalable for real-world large-scale NLP applications that require precise semantic understanding.
需要完整成稿?
PaperTan 一键生成全文 · 开题 · 降重