Thursday, April 18, 2013

Python Requests 使用摘要 二


五. 带参数访问 

1. 普通参数
>>> payload = {'key1': 'value1', 'key2': 'value2'}
>>> r = requests.get("http://httpbin.org/get", params=payload)
>>> r.url
u'http://httpbin.org/get?key2=value2&key1=value1'
>>> r.text
u'{\n  "url": "http://httpbin.org/get?key2=value2&key1=value1",\n }'

2. 带中文参数

>>> payload = {'key1': 'value1', 'key2': u'中文'}
>>> r = requests.get("http://httpbin.org/get", params=payload)
>>> r.url
u'http://httpbin.org/get?key2=%E4%B8%AD%E6%96%87&key1=value1'
>>> r.text
u'{\n  "url": "http://httpbin.org/get?key2=%E4%B8%AD%E6%96%87&key1=value1",\n }'

3. json格式
>>> import json
>>> payload = {'some': 'data'}
>>> headers = {'content-type': 'application/json'}

>>> r = requests.post(url, data=json.dumps(payload), headers=headers)


六. 文件操作

1. 文件下载
from PIL import Image
from StringIO import StringIO
i = Image.open(StringIO(r.content))
i.save('1.jpg')

2. 文件上传
>>> files = {'file': open('report.xls', 'rb')}
>>> r = requests.post(url, files=files)
>>> r.text

3. 上传时指定文件名
>>> files = {'file': ('report.xls', open('report.xls', 'rb'))}
>>> r = requests.post(url, files=files)
>>> r.text

4. 按文件接收字符串
>>> files = {'file': ('report.csv', 'some,data,to,send\nanother,row,to,send\n')}
>>> r = requests.post(url, files=files)
>>> r.text

5. 流式上传,大文件上传时,不需要全部装载到内存
with open('massive-body') as f:
    requests.post('http://some.url/streamed', data=f)

6. 流式下载
    data={'track': 'requests'}, auth=('username', 'password'), stream=True)

for line in r.iter_lines():
    if line: # filter out keep-alive new lines
        print json.loads(line)


七. 代理设置

1. 代理参数
proxies = {
  "http": "http://10.10.1.10:3128",
  "https": "http://10.10.1.10:1080",
}
requests.get("http://example.org", proxies=proxies)

>>> r = requests.get('http://ifconfig.me/ip')
>>> r.text
u'116.226.xx.xxx\n'
>>> proxies = {
...   "http": "http://175.136.xxx.xx",
... }
>>> r = requests.get("http://ifconfig.me/ip", proxies=proxies)
>>> r.text
u'175.136.xxx.xx\n'

2. 环境变量
$ export HTTP_PROXY="http://10.10.1.10:3128"
$ export HTTPS_PROXY="http://10.10.1.10:1080"
$ python
>>> import requests
>>> requests.get("http://example.org")

3. 如果代理授权方式用的是HTTP Basic Auth,则可以
proxies = {
    "http": "http://user:pass@10.10.1.10:3128/",
}


八  Session

如果需要在多次访问之间保持状态,则需要用到requests中的Session对象。
>>> s = requests.Session()
>>> r = s.get("http://httpbin.org/cookies")
>>> r.text
u'{\n  "cookies": {\n    "sessioncookie": "123456789"\n  }\n}'

另外,Session支持Keep-Alive,在同一个session中的多个请求会自动重用连接。注意,只有所有body内容读完,连接才会放回连接池重用。在使用流式文件的时候小心。


九. 异常出错

1. 网络错误(如DNS出错,链接拒绝等),抛出ConnectionError

2. 遭遇少见的非法HTTP 响应(不是requests能解析的那些404之类异常),抛出HTTPError

3. 超时,抛出Timeout

4. 301之类的跳转次数太多,抛出TooManyRedirects

5. 所有requests抛出的异常均继承自requests.exceptions.RequestException

Python Requests 使用摘要 一


Requests是由Kenneth Reitz推出的一个Python HTTP 请求操作包,在使用上,比系统自带的urllib2方便了很多,现在是1.2版本,可以通过easy_install安装。


一. 基本操作
    r = requests.get('http://httpbin.org/get')
    r = requests.post("http://httpbin.org/post")
    r = requests.put("http://httpbin.org/put")
    r = requests.delete("http://httpbin.org/delete")
    r = requests.head("http://httpbin.org/get")
    r = requests.options("http://httpbin.org/get")

    print r.headers['allow']
    HEAD, OPTIONS, GET 


二. 查看返回内容
>>> r = requests.get('http://httpbin.org/get')

1. 响应内容,以 bytes表示
>>> r.content
'{\n  "url": "http://httpbin.org/get",\n  "headers": {\n    "Content-Length": "0",\n    "Accept-Encoding": "gzip, deflate, compress",\n    "Connection": "close",\n    "Accept": "*/*",\n    "User-Agent": "python-requests/1.1.0 CPython/2.7.1 Darwin/11.4.2",\n    "Host": "httpbin.org"\n  },\n  "args": {},\n  "origin": "..."\n}'

2. 响应内容,以 unicode表示
>>> r.text              
u'{\n  "url": "http://httpbin.org/get",\n  "headers": {\n    "Content-Length": "0",\n    "Accept-Encoding": "gzip, deflate, compress",\n    "Connection": "close",\n    "Accept": "*/*",\n    "User-Agent": "python-requests/1.1.0 CPython/2.7.1 Darwin/11.4.2",\n    "Host": "httpbin.org"\n  },\n  "args": {},\n  "origin": "..."\n}'

3. 友好提示。访问无异常,为True,否则,为False。响应HTTP状态码400及以上均为False。
>>> r.ok
True

4. 访问无异常,返回空,否则,抛出异常
>>> r.raise_for_status() 

5. 响应HTTP状态码
>>> r.status_code
200

6. 访问URL
>>> r.url
u'http://httpbin.org/get'

7. 探测到headers里面的charset,要注意这里不是页面上指定的charset
>>> r.encoding

8. 如果headers里面发现不了,则会查找网页内容来探测,不过速度很慢
>>> r.apparent_encoding

9. HTTP Headers内容
>>> r.headers.keys()
['date', 'content-length', 'content-type', 'connection', 'server']
>>>r.headers['Content-Type']
'application/json'
 >>> r.headers.get('content-type')
'application/json'

10. Cookies内容
>>> r.cookies.keys()
[]
>>> cookies = dict(cookies_are='working')
>>> r = requests.get(url, cookies=cookies)
>>> r.text
u'{\n  "cookies": {\n    "cookies_are": "working"\n  }\n}'

11. 转化成json格式输出
>>> r.json()
{u'url': u'http://httpbin.org/get', u'headers': {u'Content-Length': u'0', u'Accept-Encoding': u'gzip, deflate, compress', u'Connection': u'close', u'Accept': u'*/*', u'User-Agent': u'python-requests/1.1.0 CPython/2.7.1 Darwin/11.4.2', u'Host': u'httpbin.org'}, u'args': {}, u'origin': u'...'}

12. 发出请求到接到回应的时间差
>>> r.elapsed
datetime.timedelta(0, 1, 695693)

13. iterate内容
>>> a = r.iter_lines()
>>> a
>>> a.next()
''
>>> a.next()
''
>>> a.next()
''
>>> a = r.iter_content()
>>> a
>>> a.next()
'<'
>>> a.next()
'!'
>>> a.next()
'D'
>>> a.next()
'O'
>>> a.next()
'C'

14. Timeouts
>>> requests.get('http://github.com', timeout=0.001)
Traceback (most recent call last):
  File "", line 1, in
requests.exceptions.Timeout: HTTPConnectionPool(host='github.com', port=80): Request timed out. (timeout=0.001)


三. 访问出错返回的内容
>>> r = requests.get('http://127.0.0.1:6543/abcccc')      # 这是一个不存在的链接

1. 响应内容,以 unicode表示
>>> r.text

u'<html>\n <head>\n  <title>404 Not Found</title>\n </head>\n <body>\n  <h1>404 Not Found</h1>\n  The resource could not be found.<br/><br/>\n/abcccc\n\n\n </body>\n</html>'


2. 响应内容,以 bytes表示
>>> r.content

'<html>\n <head>\n  <title>404 Not Found</title>\n </head>\n <body>\n  <h1>404 Not Found</h1>\n  The resource could not be found.<br/><br/>\n/abcccc\n\n\n </body>\n</html>'


3. 友好提示,返回False
>>> r.ok
False

4. 响应HTTP状态码
>>> r.status_code
404

5. 将访问异常外发
>>> r.raise_for_status()
Traceback (most recent call last):
  File "", line 1, in
  File "/Users/eryxlee/Workshops/python/sandbox/lib/python2.7/site-packages/requests-1.2.0-py2.7.egg/requests/models.py", line 670, in raise_for_status
    raise HTTPError(http_error_msg, response=self)
requests.exceptions.HTTPError: 404 Client Error: Not Found

6. 友好提示,返回一个合适阅读的出错提示
>>> r.reason
'Not Found'


四. 跳转情况的返回
>>> r = requests.get('http://github.com')

1. 返回内容
>>> r.status_code
200
>>> r.text[:100]
u'\n\n  http://ogp.me/ns#
fb: http://ogp.me/ns/fb# githubog: http'

2. 跳转历史
>>> r.history
(,)

3. 跳转内容
>>> r.history[0].text

u'<html>\r\n<head><title>301 Moved Permanently</title></head>\r\n<body bgcolor="white">\r\n<center><h1>301 Moved Permanently</h1></center>\r\n<hr><center>nginx</center>\r\n</body>\r\n</html>\r\n' 


4. 禁止跳转
>>> r = requests.get('http://github.com', allow_redirects=False)
>>> r.status_code
301
>>> r.history
[]

Thursday, March 28, 2013

Pyramid Route方式中减少一点add_route的方法

用了Pyramid Route方式之后,经常会面对一大堆add_route的定义,灵活利用Pyramid提供的一些便利技巧,可以大大减少这些route的定义。下面介绍一个简单的技巧:


@view_defaults(route_name='myroute' )
class MyController(object):
    def __init__(self, request):
        self.request = request
        print 'do something before every action.'

    @view_config(match_param=('ctrl=my', 'action=action1'))
    def action1(self):
        print self.request.matchdict['ctrl'], self.request.matchdict['action'], self.request.matchdict['pa']
        return Response('in my controller action 1')

    @view_config(match_param=('ctrl=my', 'action=action2'))
    def action2(self):
        print self.request.matchdict['ctrl'], self.request.matchdict['action'], self.request.matchdict['pa']
        return Response('in my controller action 2')

    @view_config(match_param=('ctrl=my', 'action=action3'), custom_predicates=(lambda context, request: request.matchdict['pa'][0]=='3',))
    def action3(self):
        print self.request.matchdict['ctrl'], self.request.matchdict['action'], self.request.matchdict['pa']
        return Response('in my controller action 3')

@view_defaults(route_name='myroute' )
class MyController2(object):
    def __init__(self, request):
        self.request = request
        print 'do something before every action.'

    @view_config(match_param=('ctrl=you','action=action1'))
    def action1(self):
        print self.request.matchdict['ctrl'], self.request.matchdict['action'], self.request.matchdict['pa']
        return Response('in you controller action 1')

    @view_config(match_param=('ctrl=you', 'action=action2'))
    def action2(self):
        print self.request.matchdict['ctrl'], self.request.matchdict['action'], self.request.matchdict['pa']
        return Response('in you controller action 2')

    @view_config(match_param=('ctrl=you', 'action=action3'), custom_predicates=(lambda context, request: request.matchdict['pa'][0]=='3',))
    def action3(self):
        print self.request.matchdict['ctrl'], self.request.matchdict['action'], self.request.matchdict['pa']
        return Response('in you controller action 3')

def main(global_config, **settings):
    """ This function returns a Pyramid WSGI application.
    """
    config = Configurator(settings=settings)
    config.add_static_view('static', 'static', cache_max_age=3600)
    config.add_route('myroute', '/{ctrl}/{action}*pa')
    config.add_route('home', '/')
    config.scan()
    return config.make_wsgi_app()

Thursday, September 6, 2012

FormEncode在SAE上的一点小改动


虽然FormEncode现在貌似用的人不多了,不过它的validate在数据校验、转换上还是很犀利的。如果用的好的话,还是可以大大减轻冗杂的代码量的。不过因为其中部分代码涉及到了一些IO操作,而SAE显然在这方面做了很多改动。因此要在SAE上使用它,就需要对它做点小小改动。主要变更代码如下:
api.py文件中有一个get_localedir函数,在系统载入时就会运行。该函数调用了resource_filename,而这个方法会去查找pwd模块(该模块SAE取消掉了)。
def get_localedir():
    """
    Retrieve the location of locales.

    If we're built as an egg, we need to find the resource within the egg.
    Otherwise, we need to look for the locales on the filesystem or in the
    system message catalog.
    """
    locale_dir = ''
    # Check the egg first
    if resource_filename is not None:
        return os.path.join(os.path.dirname(__file__), 'i18n')
        #try:
        #    locale_dir = resource_filename(__name__, "/i18n")
        #except NotImplementedError:
        #   # resource_filename doesn't work with non-egg zip files
        #    pass
    if not hasattr(os, 'access'):
        # This happens on Google App Engine
        return os.path.join(os.path.dirname(__file__), 'i18n')
    if os.access(locale_dir, os.R_OK | os.X_OK):
        # If the resource is present in the egg, use it
        return locale_dir

    # Otherwise, search the filesystem
    locale_dir = os.path.join(os.path.dirname(__file__), 'i18n')
    if not os.access(locale_dir, os.R_OK | os.X_OK):
        # Fallback on the system catalog
        locale_dir = os.path.normpath('/usr/share/locale')

    return locale_dir

大家注意一下就可以发现,这段代码里面还有个Google App Engine的代码实现。不过现在还没想到有什么好的方法来判断是否运行于SAE之上,因此,最简单的方式就是直接在行数开始处return os.path.join(os.path.dirname(__file__), 'i18n'),先当作SAE上的权宜之计,以后想到了合适的判断方式再来变更吧。

Monday, July 23, 2012

Pyramid 中SQLAlchemy 对象的JSON序列化


pyramid github中master 版本增加了custom objects的JSON支持,而在之前(1.3及以前)的版本中SQLAlchemy model对象的序列化需要编写JSONEncoder的子类,然后在dumps的时候指定。因此,想在程序中直接使用json这个renderer输出view结果是一件麻烦的事。

为了在pyramid程序中增加JSON支持,只需要在model类里面增加__json__方法即可。如

DBSession = scoped_session(sessionmaker(extension=ZopeTransactionExtension()))
Base = declarative_base()
def sqlalchemy_json(self, request):
    obj_dict = self.__dict__
    return dict((key, obj_dict[key]) for key in obj_dict if not key.startswith("_"))
Base.__json__ = sqlalchemy_json

之后继承自Base的model类即可直接json renderer中输出了。(本例中暂没测试relation)如

@view_config(route_name='home', renderer='json')
def home(request):
    one = DBSession.query(MyModel).filter(MyModel.name=='one').first()
    return {'one':one, 'project':'MyProject'}

为了支持更多第三方类的序列化,pyramid还提供了adapter的功能,如在SQLAlchemy中常用datetime数据类型,这个数据类型在序列化时也会报错,则需要增加一个adapter,如:

def datetime_adapter(obj, request):
    return obj.strftime('%Y-%m-%d %H:%M:%S')

custom_json_renderer_factory.add_adapter(datetime.datetime, datetime_adapter)

并调用config.add_renderer将 custom_json_renderer_factory注册即可。

随便附上单独提取出来的代码,直接放入项目即可在1.3版本的pyramid上使用。


import json
import datetime

from zope.interface import providedBy, Interface
from zope.interface.registry import Components

class IJSONAdapter(Interface):
    """
    Marker interface for objects that can convert an arbitrary object
    into a JSON-serializable primitive.
    """
_marker = object()

class JSON(object):
    """ Renderer that returns a JSON-encoded string.

    Configure a custom JSON renderer using the
    :meth:`~pyramid.config.Configurator.add_renderer` API at application
    startup time:

    .. code-block:: python

       from pyramid.config import Configurator

       config = Configurator()
       config.add_renderer('myjson', JSON(indent=4))

    Once this renderer is registered as above, you can use
    ``myjson`` as the ``renderer=`` parameter to ``@view_config`` or
    :meth:`~pyramid.config.Configurator.add_view``:

    .. code-block:: python

       from pyramid.view import view_config

       @view_config(renderer='myjson')
       def myview(request):
           return {'greeting':'Hello world'}

    Custom objects can be serialized using the renderer by either
    implementing the ``__json__`` magic method, or by registering
    adapters with the renderer.  See
    :ref:`json_serializing_custom_objects` for more information.

    The default serializer uses ``json.JSONEncoder``. A different
    serializer can be specified via the ``serializer`` argument.
    Custom serializers should accept the object, a callback
    ``default``, and any extra ``kw`` keyword argments passed during
    renderer construction.

    .. note::

       This feature is new in Pyramid 1.4. Prior to 1.4 there was
       no public API for supplying options to the underlying
       serializer without defining a custom renderer.
    """

    def __init__(self, serializer=json.dumps, adapters=(), **kw):
        """ Any keyword arguments will be passed to the ``serializer``
        function."""
        self.serializer = serializer
        self.kw = kw
        self.components = Components()
        for type, adapter in adapters:
            self.add_adapter(type, adapter)

    def add_adapter(self, type_or_iface, adapter):
        """ When an object of the type (or interface) ``type_or_iface`` fails
        to automatically encode using the serializer, the renderer will use
        the adapter ``adapter`` to convert it into a JSON-serializable
        object.  The adapter must accept two arguments: the object and the
        currently active request.

        .. code-block:: python

           class Foo(object):
               x = 5

           def foo_adapter(obj, request):
               return obj.x

           renderer = JSON(indent=4)
           renderer.add_adapter(Foo, foo_adapter)

        When you've done this, the JSON renderer will be able to serialize
        instances of the ``Foo`` class when they're encountered in your view
        results."""

        self.components.registerAdapter(adapter, (type_or_iface,),
            IJSONAdapter)

    def __call__(self, info):
        """ Returns a plain JSON-encoded string with content-type
        ``application/json``. The content-type may be overridden by
        setting ``request.response.content_type``."""
        def _render(value, system):
            request = system.get('request')
            if request is not None:
                response = request.response
                ct = response.content_type
                if ct == response.default_content_type:
                    response.content_type = 'application/json'
            default = self._make_default(request)
            return self.serializer(value, default=default, **self.kw)

        return _render

    def _make_default(self, request):
        def default(obj):
            if hasattr(obj, '__json__'):
                return obj.__json__(request)
            obj_iface = providedBy(obj)
            adapters = self.components.adapters
            result = adapters.lookup((obj_iface,), IJSONAdapter,
                default=_marker)
            if result is _marker:
                raise TypeError('%r is not JSON serializable' % (obj,))
            return result(obj, request)
        return default

custom_json_renderer_factory = JSON()

def datetime_adapter(obj, request):
    return obj.strftime('%Y-%m-%d %H:%M:%S')

def date_adapter(obj, request):
    return obj.strftime('%Y-%m-%d')

custom_json_renderer_factory.add_adapter(datetime.datetime, datetime_adapter)
custom_json_renderer_factory.add_adapter(datetime.date, date_adapter)

Thursday, July 5, 2012

python setuptools test command 的一点小问题


为了测试一点小程序,用pcreate -t starter sampleutils 命令生成了一个项目框架。然后在sampleutils package中建了一个tests package,然后在这个package中放置了一个test_basic.py,程序如下:

import unittest

from pyramid import testing

class View2Tests(unittest.TestCase):
    def setUp(self):
        self.config = testing.setUp()

    def tearDown(self):
        testing.tearDown()

    def test_my_view(self):
        pass

很简单,里面仅有一个testcase。现在就可以用python setup.py test来有运行单元测试了。

运行结果如下



很奇怪吧,竟然运行了两次同一个testcase。

再换nosetests运行看看:


这里明确显示是一个testcase。究竟哪里发生了问题呢?

再来看一下项目结构:


很快,我们就可以发现,pcreate生成的tests.py文件我们并没有删掉,而且这个module的名字跟tests package的名字重复了。这就是造成这次问题的原因。

但一般的,就算多了这个文件,也应该仅仅是忽略它啊,就像nosetests运行结果一样,不会去tests.py中查找任何test case。

为什么在python setup.py test 命令中会出现载入两次test_basic.py中的test case呢?

我们再来看python setup.py test这个命令的运行机制,该命令是setuptools中的一个命令之一,它通过读取setup.py文件中的配置信息,在指定目录中查找所有的testcase,然后运行。

找到setuptools的源码,打开command中的test.py程序,这个程序定义了setuptools中test command运行流程。



上面这段程序就是test case的查找过程。

从 if file.endswith('.py') and file!='__init__.py' 这一段就可以看出,它将tests.py 跟tests 混为一谈了,并且没有检查该模块是否已经载入,从而导致了在扫描tests的时候,载入了一次sampleutils.tests包,在扫描tests.py的时候,再次载入了sampleutils.tests这个包。这样也就很好的解释了在运行python setup.py test 命令的时候运行了两次相同的test case.

虽然是一个错误造成的问题,但也可以看出 setuptools中的代码还是值得推敲的。


Thursday, June 21, 2012

Mac OS Launchpad 上的重复图标


刚刚更新了一下App Store,发现Launchpad上出现了两个重复的图标iMovie,点击启动之后均打开同一个程序,版本什么的都一样,删也删不掉,好讨厌。在一番google之后,终于看到有人提出简单的修改方法,记录如下:

1. 打开终端
2. cd ~/Library/Application\ Support/Dock/
3. 这里面有一个db文件
4. 做个备份
5. 用sqlite3打开它
6. 查找其中的apps表,select * from apps where title='iMovie';
7. 发现果然有重复的,删掉id大的条目
8. killall Dock
9. 正常了

貌似还有个叫launchpad-control的工具可以完成Launchpad的管理,通过http://chaosspace.de/launchpad-control/下载(墙外)。

环境:Mac OS X 10.7.4