推荐学习书目
› Learn Python the Hard Way
Python Sites
› PyPI - Python Package Index
› http://diveintopython.org/toc/index.html
› Pocoo
值得关注的项目
› PyPy
› Celery
› Jinja2
› Read the Docs
› gevent
› pyenv
› virtualenv
› Stackless Python
› Beautiful Soup
› 结巴中文分词
› Green Unicorn
› Sentry
› Shovel
› Pyflakes
› pytest
Python 编程
› pep8 Checker
Styles
› PEP 8
› Google Python Style Guide
› Code Style from The Hitchhiker's Guide
zeroten
V2EX  ›  Python

python2 如何正确的处理 4 字节的字符,为什么一个字符变成了两个?

  •  
  •   zeroten ·
    ladder1984 · Jan 21, 2016 · 3674 views
    This topic created in 3913 days ago, the information mentioned may be changed or developed.
    x = u'\U0001f604abc'
    
    print('length:',len(x))
    for i in x:
        print(i)
    

    得到输出:

    ('length:', 5)
    �
    �
    a
    b
    c

    x 是 4 个字符,其中第一个是 4 字节字符,一个笑脸表情的 unicdoe 码,现在显然被拆分成了两个。我写的过滤函数就过滤失败了:

    def filter_invalid_str(text):
    return ''.join(map(lambda x: x if u'\u0000' < x < u'\uFFFF' else '_', text))

    所以,明明一个字符为什么变成了两个,如何当作一个字符处理?

    8 replies  •  2016-01-21 20:07:16 +08:00
    Zzzzzzzzz
        1
    Zzzzzzzzz  
       Jan 21, 2016
    编译 python2 的时候加上--enable-unicode=ucs4
    Hackathon
        2
    Hackathon  
       Jan 21, 2016
    zeroten
        3
    zeroten  
    OP
       Jan 21, 2016
    或者,过滤掉 u'\u0000' < x < u'\uFFFF'之外的字符有什么好的实现方法?
    zeroten
        4
    zeroten  
    OP
       Jan 21, 2016
    @Zzzzzzzzz 这个编译是指?编译 python 的 C 源码么?
    Zzzzzzzzz
        7
    Zzzzzzzzz  
       Jan 21, 2016
    @zeroten 对, 你现在是 ucs2.
    zeroten
        8
    zeroten  
    OP
       Jan 21, 2016
    @Hackathon
    @Zzzzzzzzz

    很好奇为什么落在[\uD800-\uDBFF][\uDC00-\uDFFF]这个范围内
    About   ·   Help   ·   Advertise   ·   Blog   ·   API   ·   FAQ   ·   Privacy   ·   Solana   ·   917 Online   Highest 6679   ·     Select Language
    创意工作者们的社区
    World is powered by solitude
    VERSION: 3.9.8.5 · 30ms · UTC 20:16 · PVG 04:16 · LAX 13:16 · JFK 16:16
    ♥ Do have faith in what you're doing.