python使用正则表达式去除中文文本多余空格，保留英文之间空格方法详解_随笔

python使用正则表达式去除中文文本多余空格，保留英文之间空格方法详解

在pdf转为文本的时候，经常会多出空格，影响数据观感，因此需要去掉文本中多余的空格，而文本中的英文之间的正常空格需要保留，输入输出如下：

input：我今天赚了 10 个亿，老百姓very happy。

output：我今天赚了10个亿，老百姓very happy。

代码

def clean_space(text):
  """"
  处理多余的空格
  """
  match_regex = re.compile(u'[u4e00-u9fa5。.,，:：《》、()（）]{1} +(?

python去除英文单词之间多余的空格
re.sub(" +", " ", s)

import re 

s = "     info has been found (+/- 100 pages, and 4.5 mb of .pdf files) now i have to wait untill our team leader has processed it and learns html.     "
re.sub(" +", " ", s)

' '.join(s.split())

s = "     info has been found (+/- 100 pages, and 4.5 mb of .pdf files) now i have to wait untill our team leader has processed it and learns html.     "

s = ' '.join(s.split())
s

更多关于python使用正则表达式去除多余空格方法请查看下面的相关链接					
										


					
						欢迎分享，转载请注明来源：内存溢出
原文地址: http://outofmemory.cn/zaji/3238331.html

python使用正则表达式去除中文文本多余空格，保留英文之间空格方法详解

发表评论

评论列表（0条）