在Python3 Pandas中读取/导入CSV文件时出现编码问题
作者:互联网
我正在尝试使用熊猫读取电影镜头数据集:http://files.grouplens.org/datasets/movielens/ml-100k/.
我正在使用Python 3.4版,并且正在按照“ http://www.gregreda.com/2013/10/26/using-pandas-on-the-movielens-dataset/”中给出的教程进行操作.
当我尝试使用此处提到的代码读取u.item数据时:
# the movies file contains columns indicating the movie's genres
# let's only load the first five columns of the file with usecols
m_cols = ['movie_id', 'title', 'release_date', 'video_release_date', 'imdb_url']
movies = pd.read_csv('ml-100k/u.item', sep='|', names=m_cols, usecols=range(5), encoding='UTF-8')
我收到以下错误“ UnicodeDecodeError:’utf-8’编解码器无法解码位置3的字节0xe9:无效的连续字节”.
什么可能是此错误的可能原因,以及什么解决方案
我尝试在pd.read_csv(encoding =’utf-8′)中添加encoding =’utf-8′,但是不幸的是它并没有解决任何问题.
错误回溯为:
---------------------------------------------------------------------------
UnicodeDecodeError Traceback (most recent call last)
<ipython-input-4-4cc01a7faf02> in <module>()
9 # let's only load the first five columns of the file with usecols
10 m_cols = ['movie_id', 'title', 'release_date', 'video_release_date', 'imdb_url']
---> 11 movies = pd.read_csv('ml-100k/u.item', sep='|', names=m_cols, usecols=range(5), encoding='UTF-8')
/usr/local/lib/python3.4/site-packages/pandas/io/parsers.py in parser_f(filepath_or_buffer, sep, dialect, compression, doublequote, escapechar, quotechar, quoting, skipinitialspace, lineterminator, header, index_col, names, prefix, skiprows, skipfooter, skip_footer, na_values, na_fvalues, true_values, false_values, delimiter, converters, dtype, usecols, engine, delim_whitespace, as_recarray, na_filter, compact_ints, use_unsigned, low_memory, buffer_lines, warn_bad_lines, error_bad_lines, keep_default_na, thousands, comment, decimal, parse_dates, keep_date_col, dayfirst, date_parser, memory_map, float_precision, nrows, iterator, chunksize, verbose, encoding, squeeze, mangle_dupe_cols, tupleize_cols, infer_datetime_format, skip_blank_lines)
472 skip_blank_lines=skip_blank_lines)
473
--> 474 return _read(filepath_or_buffer, kwds)
475
476 parser_f.__name__ = name
/usr/local/lib/python3.4/site-packages/pandas/io/parsers.py in _read(filepath_or_buffer, kwds)
258 return parser
259
--> 260 return parser.read()
261
262 _parser_defaults = {
/usr/local/lib/python3.4/site-packages/pandas/io/parsers.py in read(self, nrows)
719 raise ValueError('skip_footer not supported for iteration')
720
--> 721 ret = self._engine.read(nrows)
722
723 if self.options.get('as_recarray'):
/usr/local/lib/python3.4/site-packages/pandas/io/parsers.py in read(self, nrows)
1168
1169 try:
-> 1170 data = self._reader.read(nrows)
1171 except StopIteration:
1172 if nrows is None:
pandas/parser.pyx in pandas.parser.TextReader.read (pandas/parser.c:7544)()
pandas/parser.pyx in pandas.parser.TextReader._read_low_memory (pandas/parser.c:7784)()
pandas/parser.pyx in pandas.parser.TextReader._read_rows (pandas/parser.c:8617)()
pandas/parser.pyx in pandas.parser.TextReader._convert_column_data (pandas/parser.c:9928)()
pandas/parser.pyx in pandas.parser.TextReader._convert_tokens (pandas/parser.c:10714)()
pandas/parser.pyx in pandas.parser.TextReader._convert_with_dtype (pandas/parser.c:12118)()
pandas/parser.pyx in pandas.parser.TextReader._string_convert (pandas/parser.c:12283)()
pandas/parser.pyx in pandas.parser._string_box_utf8 (pandas/parser.c:17655)()
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 3: invalid continuation byte
解决方法:
如果找到解决问题的两种可能的技巧:
1 /在文本编辑器中打开文件,并保存编码为“ UTF-8”的文件
===>例如,在Sublime Text中,遵循以下标签:>>>“ Edit”>>>“ Save with encoding”>> “ UTF-8”
2 /或仅使用python2打开文档…
找不到更好的解决方案.
标签:data-analysis,pandas,python-3-x,csv,python 来源: https://codeday.me/bug/20191120/2042013.html