
   [1]SourceForge.net Logo 

                                  pullparser

   A simple "pull API" for HTML parsing, after Perl's HTML::TokeParser.
   Many simple HTML parsing tasks are simpler this way than with the
   HTMLParser module. pullparser.PullParser is a subclass of
   HTMLParser.HTMLParser.

   Examples:

   This program extracts all links from a document. It will print one
   line for each link, containing the URL and the textual description
   between the <a>...</a> tags:
import pullparser, sys
f = file(sys.argv[1])
p = pullparser.PullParser(f)
for token in p.tags("a"):
    if token.type == "endtag": continue
    url = dict(token.attrs).get("href", "-")
    text = p.get_compressed_text(endat=("endtag", "a"))
    print "%s\t%s" % (url, text)

   This program extracts the <title> from the document:
import pullparser, sys
f = file(sys.argv[1])
p = pullparser.PullParser(f)
if p.get_tag("title"):
    title = p.get_compressed_text()
    print "Title: %s" % title

   Thanks to Gisle Aas, who wrote HTML::TokeParser.

Download

   All documentation (including this web page) is included in the
   distribution.

   Development release.
     * [2]pullparser-0.0.5b.tar.gz
     * [3]pullparser-0_0_5b.zip
     * [4]Change Log (included in distribution)
     * [5]Older versions.

   For installation instructions, see the INSTALL file included in the
   distribution.

FAQs

     * Which version of Python do I need?
       2.2 or above.
     * Which license?
       The [6]Perl Artistic license (included in distribution). This may
       change to BSD (more liberal) at some point.
     * Why don't I see the tokens I expect?
          + Are there missing end-tags in your HTML? (Maybe this will
            improve in future.)
          + Element names passed to methods such as
            PullParser.get_token() must be given in lower case - maybe
            you forgot that? (Element names in the HTML can be any case,
            of course.)

   [7]John J. Lee, May 2004.
     _________________________________________________________________

   [8]Home
   [9]ClientCookie
   [10]ClientCookie docs
   [11]ClientForm
   [12]DOMForm
   [13]python-spidermonkey
   [14]ClientTable
   [15]mechanize
   pullparser
   [16]General FAQs
   [17]1.5.2 urllib2.py
   [18]1.5.2 urllib.py
   [19]Download
   [20]FAQs

References

   1. http://sourceforge.net/
   2. http://wwwsearch.sourceforge.net/pullparser/src/pullparser-0.0.5b.tar.gz
   3. http://wwwsearch.sourceforge.net/pullparser/src/pullparser-0_0_5b.zip
   4. http://wwwsearch.sourceforge.net/pullparser/src/ChangeLog.txt
   5. http://wwwsearch.sourceforge.net/pullparser/src/
   6. http://www.opensource.org/licenses/artistic-license.php
   7. mailto:jjl@pobox.com
   8. http://wwwsearch.sourceforge.net/
   9. http://wwwsearch.sourceforge.net/ClientCookie/
  10. http://wwwsearch.sourceforge.net/ClientCookie/doc.html
  11. http://wwwsearch.sourceforge.net/ClientForm/
  12. http://wwwsearch.sourceforge.net/DOMForm/
  13. http://wwwsearch.sourceforge.net/python-spidermonkey/
  14. http://wwwsearch.sourceforge.net/ClientTable/
  15. http://wwwsearch.sourceforge.net/mechanize/
  16. http://wwwsearch.sourceforge.net/bits/clientx.html
  17. http://wwwsearch.sourceforge.net/bits/urllib2_152.py
  18. http://wwwsearch.sourceforge.net/bits/urllib_152.py
  19. http://wwwsearch.sourceforge.net/pullparser/#download
  20. http://wwwsearch.sourceforge.net/pullparser/#faq
