经过预训练的Fasttext模型返回词汇表外单词的胡言乱语

2024-10-04 11:25:53 发布

您现在位置:Python中文网/ 问答频道 /正文

我在使用预先培训过的快速文本.bin模型(检索自https://fasttext.cc/docs/en/crawl-vectors.html)。检查词汇表中最相似的单词会返回合理的响应。然而,当检查单词表中只有一个字符不同的单词最相似时,会返回胡言乱语。在

我的问题是:这与模型有关,还是我用错了方法?在

from gensim.models.wrappers import FastText
model = FastText.load_fasttext_format('cc.en.300.bin')
model.most_similar("universitet")
[('Universitet', 0.8522759675979614),
 ('högskolan', 0.677900493144989),
 ('Högskola', 0.6725144386291504),
 ('högskola', 0.6724666357040405),
 ('Högskolan', 0.6600401997566223),
 ('Universitetet', 0.6519213318824768),
 ('Høgskolen', 0.647462010383606),
 ('Universiteti', 0.6399329900741577),
 ('forskning', 0.617483377456665),
 ('språk', 0.6172543168067932)]

model.most_similar("universitett")
[('ESTATERETAILCONSUMERPHONESCARSBIKESAPPSINTERNETTABLETSCOMPUTERSSOCIETYPOLITICSLAWCRIMEENVIRONMENTSCIENCEARTSCELEBRITIESSPORTSSPECIALSFIRST',
  0.47905537486076355),
 ('Wikipedia-Page-Suzannah-B-Troy-6-yrs-after-Misogynist-Cyber-Vandalism-Censorship-via-Deletion-on-a-page-about-Censorship-Wikipedia-Agrees-to-retur',
  0.47733378410339355),
 ('DEky4M0BSpUOTPnSpkuL5I0GTSnRI4jMepcaFAoxIoFnX5kmJQk1aYvr2odGBAAIfkECQoABAAsCQAAABAAEgAACGcAARAYSLCgQQEABBokkFAhAQEQHQ4EMKCiQogRCVKsOOAiRocbLQ7EmJEhR4cfEWoUOTFhRIUNE44kGZOjSIQfG9rsyDCnzp0AaMYMyfNjS6JFZWpEKlDiUqALJ0KNatKmU4NDBwYEACH5BAkKAAQALAkAAAAQABIAAAhpAAEQGEiQIICDBAUgLEgAwICHAgkImBhxoMOHAyJOpGgQY8aBGxV2hJgwZMWLFTcCUIjwoEuLBym69PgxJMuDNAUqVDkz50qZLi',
  0.474983274936676),
 ('DEky4M0BSpUOTPnSpkuL5I0GTSnRI4jMepcaFAoxIoFnX5kmJQk1aYvr2odGBAAIfkECQoABAAsCQAAABAAEgAACGcAARAYSLCgQQEABBokkFAhAQEQHQ4EMKCiQogRCVKsOOAiRocbLQ7EmJEhR4cfEWoUOTFhRIUNE44kGZOjSIQfG9rsyDCnzp0AaMYMyfNjS6JFZWpEKlDiUqALJ0KNatKmU4NDBwYEACH5BAUKAAQALAkAAAAQABIAAAhpAAEQGEiQIICDBAUgLEgAwICHAgkImBhxoMOHAyJOpGgQY8aBGxV2hJgwZMWLFTcCUIjwoEuLBym69PgxJMuDNAUqVDkz50qZLi',
  0.47364047169685364),
 ('crescendosexibloguerobateyabsorbersexiindesignabledinerolatifundiosexibrezarcularsutesexirapoplinbrezarcorrentosoVd.lazadareflejoreglafeministabrezarchuzasexiouttiqueblogueroin',
  0.47090965509414673),
 ('QQFZAAEACwAAAAAGQASAAAIjgAJCBQIoGDBgQgTKiwooGHDgwshDgTgsOLDhAAGaAQwUYBBhx85EtS4cWLGjR5JSjxZkgDFkwwLohTJUqTLlANiwvQ4seVNjwwfBoVokKjFo0Jlksz506NFiklZtoQKFSjIoktLVv1YsahSn1WP0vzq02VYoAjJMsVYVKHZrDbdupW6Vq5cunHtRjQoMCAAIfkECRQABAAsCQADAAQABAAACAsABQgkILCgwYEBAQAh',
  0.46747487783432007),
 ('записиТелепрограммаVikerraadioOtseEsilehtJärelkuulamineSaatekavaPodcastidRaadioteaterRaadio',
  0.4659830331802368),
 ('deblogueroreflejoantecedentesexitlacuachebateysuteindesignableabsorbersexilatifundiosexibrezarsutemultiétnicosexiplinrapobrezarcorrentosoVd.lazadafisiochillidomabrezarsico-chuzaoutcolodrablogueroin',
  0.46159273386001587),
 ('2OtseEsilehtJärelkuulamineSaatedPodcastidKlassikaraadioOtseEsilehtJärelkuulamineSaatekavaPodcastidRaadio',
  0.4609595537185669),
 ('leilighetEiendomstypeSelveierleilighetPlass', 0.4550461769104004)]

Tags: https模型文本mostmodelbinwikipedia单词