基于python生成词云图的代码示例

更新时间：2023年11月16日 08:32:07 作者：颜酱

这篇文章主要个介绍了如何基于python生成词云图的代码示例,文中有详细的代码示例喝图文讲解,对大家的学习或工作有一定的帮助,需要的朋友可以参考下

1. 读文件

使用了下 codecs 库，读取内容更方便。

import codecs
def get_file_content(filePath):
  with codecs.open(filePath, 'r', 'utf-8') as f:
      txt = f.read()
  return txt

2. 分词

jieba是一个分词库，可以将一段文本分割成词语。cut 是将文本精确切分开，不存在冗余词语。比如颜酱是一个厉害的厨师，会变成['颜酱', '厉害', '厨师']。 Counter是一个计数器，可以统计词语出现的次数，most_common 是取出最常用的词语。

import jieba
from collections import Counter
def get_words(txt):
    # 先分词，得到词语数组
    seg_list = jieba.cut(txt)
    # 开始计数
    c = Counter()
    for x in seg_list:
        if len(x)>1 and x != '\r\n':
            c[x] += 1
    word_list = []
    print('常用词频度统计结果')
     # 统计前99个词
    for (k,v) in c.most_common(99):
        word_list.append(str(k))
    # 将词语生成文本文件
    file = open("./dist/out_words.txt", 'w').close()
    with open("./dist/out_words.txt",'a+',encoding='utf-8') as writeFile:
        for (k,v) in c.most_common(99):    # 统计前99个词
            writeFile.write(str(k))
            writeFile.write(str(v))
            writeFile.write('\n')
    print(word_list)
    # ['发展','平安']
    return word_list

3. 生成云图

wordcloud 是一个词云库，可以将词语生成词云图片。

import wordcloud;
def generate_cloud_image(file_path, shape_image_path):
  word_list = get_words(get_file_content(file_path))
  string = ' '.join(word_list)
  # 读取词云形状图片
  image = imageio.v2.imread(shape_image_path)
  # 先实例化一个词云对象
  wc = wordcloud.WordCloud(width=image.shape[0],     # 词云图宽度同原图片宽度
                          height=image.shape[1],
                          background_color='white',  # 背景颜色白色
                          font_path='Arial Unicode.ttf',    # 指定字体路径，微软雅黑，可从自带的字体库中找
                          mask=image,   # mask 指定词云形状图片，默认为矩形
                          scale=3)      # 默认为1，越大越清晰
  # 生成词云
  wc.generate(string)
  # 保存成文件,output_wordcloud.png，词云图
  wc.to_file('dist/output_wordcloud.png')
  # 弹出图片显示
  alert_image('dist/output_wordcloud.png')

4. 显示词云图

matplotlib是一个绘图库，可以将图片显示出来，plt 用于显示图片，mpimg 用于读取图片。

#
# plt用于显示图片
import matplotlib.pyplot as plt
# mpimg 用于读取图片
import matplotlib.image as mpimg
def alert_image(image_path):
  # 这里是让词云图弹出来显示
  lena = mpimg.imread(image_path) # 读取和代码处于同一目录下的 lena.png
  # 此时 lena 就已经是一个 np.array 了，可以对它进行任意处理
  lena.shape #(512, 512, 3)
  plt.imshow(lena) # 显示图片
  plt.axis('off') # 不显示坐标轴
  plt.show()

这就可以了！

其他文件

准备好文件实验：

input.txt

cloud.jpg

generate_cloud_image.py

generate_cloud_image.py:

import codecs
import jieba
import imageio
import wordcloud
import matplotlib.pyplot as plt
import matplotlib.image as mpimg
from collections import Counter
# def get_words(txt):...
# def get_file_content:...
# def alert_image(image_path):...
# def generate_cloud_image(file_path, shape_image_path):...
generate_cloud_image('input.txt', 'cloud.jpg')

完整版：

import codecs
import jieba
import imageio
import wordcloud
import matplotlib.pyplot as plt
import matplotlib.image as mpimg
from collections import Counter
#  get_words函数用于统计词频，生成out.txt，展示词语和词频 如发展218 坚持170
def get_words(txt):
    seg_list = jieba.cut(txt)
    c = Counter()
    for x in seg_list:
        if len(x)>1 and x != '\r\n':
            c[x] += 1
    word_list = []
    print('常用词频度统计结果')
     # 统计前99个词
    for (k,v) in c.most_common(99):
        word_list.append(str(k))

    file = open("./out_words.txt", 'w').close()
    with open("./out_words.txt",'a+',encoding='utf-8') as writeFile:
        for (k,v) in c.most_common(99):    # 统计前99个词
            writeFile.write(str(k))
            writeFile.write(str(v))
            writeFile.write('\n')
    print(word_list)
    # ['发展','平安']
    return word_list

def get_file_content(filePath):
    with codecs.open(filePath, 'r', 'utf-8') as f:
        txt = f.read()
    return txt
def alert_image(image_path):
    # 这里是让词云图弹出来显示
    lena = mpimg.imread(image_path) # 读取和代码处于同一目录下的 lena.png
    # 此时 lena 就已经是一个 np.array 了，可以对它进行任意处理
    lena.shape #(512, 512, 3)
    plt.imshow(lena) # 显示图片
    plt.axis('off') # 不显示坐标轴
    plt.show()

# 根据文件生成词云图，file_path是文本文件路径，shape_image_path是词云图的图片途径
def generate_cloud_image(file_path, shape_image_path):
    word_list = get_words(get_file_content(file_path))
    string = ' '.join(word_list)
    # 读取词云形状图片
    image = imageio.v2.imread(shape_image_path)
    # 生成词云图片，先实例化一个词云对象
    wc = wordcloud.WordCloud(width=image.shape[0],     # 词云图宽度同原图片宽度
                            height=image.shape[1],
                            background_color='white',  # 背景颜色白色
                            font_path='Arial Unicode.ttf',    # 指定字体路径，微软雅黑，可从win自带的字体库中找
                            mask=image,   # mask 指定词云形状图片，默认为矩形
                            scale=3)      # 默认为1，越大越清晰
    # 再给词云
    wc.generate(string)
    # 保存成文件,output_wordcloud.png，词云图
    wc.to_file('output_wordcloud.png')
    alert_image('output_wordcloud.png')

generate_cloud_image('input.txt', 'cloud.jpg')

运行

记得先pip3 install jieba imageio wordcloud matplotlib然后python3 generate_cloud_image.py📢：文件在同一目录，进到这个目录下运行命令

以上就是基于python生成词云图的代码示例的详细内容，更多关于python生成词云图的资料请关注脚本之家其它相关文章！

您可能感兴趣的文章:

Python3.7 + Yolo3实现识别语音播报功能
这篇文章主要介绍了Python3.7 + Yolo3识别语音播报功能,开始之前我们先得解析出来Yolo3的代码，从而获取到被识别出来的物体标签，具体详细过程跟随小编一起看看吧
2021-12-12
详解Python中break语句的用法
这篇文章主要介绍了详解Python中break语句的用法,是Python入门的呼出知识,需要的朋友可以参考下
2015-05-05
python PyAUtoGUI库实现自动化控制鼠标键盘
这篇文章主要介绍了python PyAUtoGUI库实现自动化控制鼠标键盘，帮助大家更好的理解和使用python，感兴趣的朋友可以了解下
2020-09-09
Python编写车票订购系统 Python实现快递收费系统
这篇文章主要为大家详细介绍了Python编写车票订购系统，Python实现快递收费系统，文中示例代码介绍的非常详细，具有一定的参考价值，感兴趣的小伙伴们可以参考一下
2022-08-08
python os.path.isfile 的使用误区详解
今天小编就为大家分享一篇python os.path.isfile 的使用误区详解，具有很好的参考价值，希望对大家有所帮助。一起跟随小编过来看看吧
2019-11-11
python通过urllib2获取带有中文参数url内容的方法
这篇文章主要介绍了python通过urllib2获取带有中文参数url内容的方法,涉及Python中文编码的技巧,具有一定参考借鉴价值,需要的朋友可以参考下
2015-03-03
Python变量访问权限控制详解
这篇文章主要介绍了Python变量访问权限控制详解,文中通过示例代码介绍的非常详细，对大家的学习或者工作具有一定的参考学习价值,需要的朋友可以参考下
2019-06-06
Python中常用信号signal类型实例
这篇文章主要介绍了Python中常用信号signal类型实例，分享了相关代码示例，小编觉得还是挺不错的，具有一定借鉴价值，需要的朋友可以参考下
2018-01-01
Python元字符的用法实例解析
这篇文章主要介绍了Python元字符的用法实例解析，具有一定借鉴价值,需要的朋友可以参考下
2018-01-01
Python自动化测试框架pytest的详解安装与运行
这篇文章主要为大家介绍了Python自动化测试框架pytest的简介以及安装与运行，有需要的朋友可以借鉴参考下希望能够有所帮助，祝大家多多进步
2021-10-10