一、安装Spark
- 检查基础环境hadoop,jdk
- 下载spark
- 解压,文件夹重命名、权限
- 配置文件
- 环境变量
- 试运行Python代码
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220301085500668-427224852.png)
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220301092641185-229826461.png)
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220301092624728-852872443.png)
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220301085526577-1436993924.png)
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220301085920742-2016622183.png)
二、Python编程练习:英文文本的词频统计
- 准备文本文件
- 读文件
- 预处理:大小写,标点符号,停用词
- 分词
- 统计每个单词出现的次数
- 按词频大小排序
- 结果写文件
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220305134421051-1074789754.png)
读文件:
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220305141812487-1859954854.png)
预处理:大小写,标点符号,停用词、分词
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220305141858189-1315570384.png)
统计每个单词出现的次数,按词频大小排序
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220305142022671-1970785713.png)
结果写文件
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220305142054192-1073127156.png)
输出
![](https://www.icode9.com/i/l/?n=22&i=blog/2764534/202203/2764534-20220305142125108-2030049396.png)
标签:文件,Python,练习,词频,大小写,Spark,分词
来源: https://www.cnblogs.com/rromi/p/15949975.html