{
  "attachments": [],
  "comments_archived": true,
  "date": "2003-08-23T18:52:14.000Z",
  "layout": "post",
  "title": "Scraping with web services: Success",
  "wordpress_id": 467,
  "wordpress_slug": "rss-scrape-urls2",
  "wordpress_url": "http://www.decafbad.com/blog/?p=467",
  "year": "2003",
  "month": "08",
  "day": "23",
  "isDir": false,
  "slug": "rss-scrape-urls2",
  "type": "entry",
  "postName": "2003-08-23-rss-scrape-urls2",
  "html": "<p>Okay, so I took another shot at <a href=\"http://www.decafbad.com/blog/geek/rss_scrape_urls.html\">scraping <span class=\"caps\">HTML</span> with web services</a> with <a href=\"http://www.jlist.com\">another site</a> that passes the <span class=\"caps\">HTML </span>Tidy step.  Luckily, this is a site that I already scrape using my own tool, so I have XPath expressions already cooked up to dig out info for <span class=\"caps\">RSS</span> items.  So, here are the vitals:</p>\n\n\n<ul>\n    <li>Site: <a href=\"http://www.jlist.com\">http://www.jlist.com</a></li>\n    <li><span class=\"caps\">XSL</span>: <a href=\"http://www.decafbad.com/jlist.xsl\">http://www.decafbad.com/jlist.xsl</a></li>\n    <li>Tidy <span class=\"caps\">URL</span>: <a href=\"http://cgi.w3.org/cgi-bin/tidy?docAddr=http%3A%2F%2Fwww.jlist.com%2FUPDATES%2FPG%2F365%2F\">http://cgi.w3.org/cgi-bin/tidy?\n\n</a><p><a href=\"http://cgi.w3.org/cgi-bin/tidy?docAddr=http%3A%2F%2Fwww.jlist.com%2FUPDATES%2FPG%2F365%2F\">docAddr=http%3A%2F%2F</a><a href=\"http://www.jlist.com%2FUPDATES%2FPG%2F365%2F\">www.jlist.com%2FUPDATES%2FPG%2F365%2F</a></p></li>\n    <li>Final <span class=\"caps\">URL</span>: <a href=\"http://www.w3.org/2000/06/webdata/xslt?xslfile=http%3A%2F%2Fwww.decafbad.com%2Fjlist.xsl&amp;xmlfile=http%3A%2F%2Fcgi.w3.org%2Fcgi-bin%2Ftidy%3FdocAddr%3Dhttp%253A%252F%252Fwww.jlist.com%252FUPDATES%252FPG%252F365%252F&amp;transform=Submit\">http://www.w3.org/2000/06/webdata/xslt?<p></p>\n<p>xslfile=http%3A%2F%2Fwww.decafbad.com%2Fjlist.xsl&amp;</p>\n<p>xmlfile=http%3A%2F%2Fcgi.w3.org%2Fcgi-bin%2Ftidy%3F</p>\n<p>docAddr%3Dhttp%253A%252F%252Fwww.jlist.com%252FUPDATES%252FPG%252F365%252F&amp;</p>\n</a><p><a href=\"http://www.w3.org/2000/06/webdata/xslt?xslfile=http%3A%2F%2Fwww.decafbad.com%2Fjlist.xsl&amp;xmlfile=http%3A%2F%2Fcgi.w3.org%2Fcgi-bin%2Ftidy%3FdocAddr%3Dhttp%253A%252F%252Fwww.jlist.com%252FUPDATES%252FPG%252F365%252F&amp;transform=Submit\">transform=Submit</a></p></li><p></p>\n</ul>\n\n\n\n<pre><code>&lt;p&gt;Unfortunately, although it looks okay to me, this feed &lt;a href=\"http://feeds.archive.org/validator/check?url=http%3A%2F%2Fwww.w3.org%2F2000%2F06%2Fwebdata%2Fxslt%3Fxslfile%3Dhttp%253A%252F%252Fwww.decafbad.com%252Fjlist.xsl%26xmlfile%3Dhttp%253A%252F%252Fcgi.w3.org%252Fcgi-bin%252Ftidy%253FdocAddr%253Dhttp%25253A%25252F%25252Fwww.jlist.com%25252FUPDATES%25252FPG%25252F365%25252F%26transform%3DSubmit\"&gt;doesn&amp;#8217;t validate yet&lt;/a&gt;, but I&amp;#8217;m still poking around with it to get things straight.  Feel free to help me out!  :)&lt;/p&gt;\n</code></pre>\n<!--more-->\n\n\n<p>shortname=rss_scrape_urls2</p>\n",
  "body": "<p>Okay, so I took another shot at <a href=\"http://www.decafbad.com/blog/geek/rss_scrape_urls.html\">scraping <span class=\"caps\">HTML</span> with web services</a> with <a href=\"http://www.jlist.com\">another site</a> that passes the <span class=\"caps\">HTML </span>Tidy step.  Luckily, this is a site that I already scrape using my own tool, so I have XPath expressions already cooked up to dig out info for <span class=\"caps\">RSS</span> items.  So, here are the vitals:</p>\r\n\r\n\r\n<ul>\r\n\t<li>Site: <a href=\"http://www.jlist.com\">http://www.jlist.com</a></li>\r\n\t<li><span class=\"caps\">XSL</span>: <a href=\"http://www.decafbad.com/jlist.xsl\">http://www.decafbad.com/jlist.xsl</a></li>\r\n\t<li>Tidy <span class=\"caps\">URL</span>: <a href=\"http://cgi.w3.org/cgi-bin/tidy?docAddr=http%3A%2F%2Fwww.jlist.com%2FUPDATES%2FPG%2F365%2F\">http://cgi.w3.org/cgi-bin/tidy?<br />docAddr=http%3A%2F%2Fwww.jlist.com%2FUPDATES%2FPG%2F365%2F</a></li>\r\n\t<li>Final <span class=\"caps\">URL</span>: <a href=\"http://www.w3.org/2000/06/webdata/xslt?xslfile=http%3A%2F%2Fwww.decafbad.com%2Fjlist.xsl&#38;xmlfile=http%3A%2F%2Fcgi.w3.org%2Fcgi-bin%2Ftidy%3FdocAddr%3Dhttp%253A%252F%252Fwww.jlist.com%252FUPDATES%252FPG%252F365%252F&#38;transform=Submit\">http://www.w3.org/2000/06/webdata/xslt?<br />xslfile=http%3A%2F%2Fwww.decafbad.com%2Fjlist.xsl&#38;<br />xmlfile=http%3A%2F%2Fcgi.w3.org%2Fcgi-bin%2Ftidy%3F<br />docAddr%3Dhttp%253A%252F%252Fwww.jlist.com%252FUPDATES%252FPG%252F365%252F&#38;<br />transform=Submit</a></li>\r\n</ul>\r\n\r\n\t<p>Unfortunately, although it looks okay to me, this feed <a href=\"http://feeds.archive.org/validator/check?url=http%3A%2F%2Fwww.w3.org%2F2000%2F06%2Fwebdata%2Fxslt%3Fxslfile%3Dhttp%253A%252F%252Fwww.decafbad.com%252Fjlist.xsl%26xmlfile%3Dhttp%253A%252F%252Fcgi.w3.org%252Fcgi-bin%252Ftidy%253FdocAddr%253Dhttp%25253A%25252F%25252Fwww.jlist.com%25252FUPDATES%25252FPG%25252F365%25252F%26transform%3DSubmit\">doesn&#8217;t validate yet</a>, but I&#8217;m still poking around with it to get things straight.  Feel free to help me out!  :)</p>\r\n<!--more-->\r\nshortname=rss_scrape_urls2\r\n",
  "parentPath": "./content/posts/archives/2003",
  "path": "2003/08/23/rss-scrape-urls2",
  "summary": "Okay, so I took another shot at scraping HTML with web services with another site that passes the HTML Tidy step.  Luckily, this is a site that I already scrape using my own tool, so I have XPath expressions already cooked up to dig out info for RSS items.  So, here are the vitals:\n\n\n\n    Site: http://www.jlist.com\n    XSL: http://www.decafbad.com/jlist.xsl\n    Tidy URL: http://cgi.w3.org/cgi-bin/tidy?\n\ndocAddr=http%3A%2F%2Fwww.jlist.com%2FUPDATES%2FPG%2F365%2F\n    Final URL: http://www.w3.org/2000/06/webdata/xslt?\nxslfile=http%3A%2F%2Fwww.decafbad.com%2Fjlist.xsl&\nxmlfile=http%3A%2F%2Fcgi.w3.org%2Fcgi-bin%2Ftidy%3F\ndocAddr%3Dhttp%253A%252F%252Fwww.jlist.com%252FUPDATES%252FPG%252F365%252F&\ntransform=Submit\n\n\n\n\n<p>Unfortunately, although it looks okay to me, this feed <a href=\"http://feeds.archive.org/validator/check?url=http%3A%2F%2Fwww.w3.org%2F2000%2F06%2Fwebdata%2Fxslt%3Fxslfile%3Dhttp%253A%252F%252Fwww.decafbad.com%252Fjlist.xsl%26xmlfile%3Dhttp%253A%252F%252Fcgi.w3.org%252Fcgi-bin%252Ftidy%253FdocAddr%253Dhttp%25253A%25252F%25252Fwww.jlist.com%25252FUPDATES%25252FPG%25252F365%25252F%26transform%3DSubmit\">doesn&#8217;t validate yet</a>, but I&#8217;m still poking around with it to get things straight.  Feel free to help me out!  :)</p>",
  "needsBuild": true,
  "prevPostPath": "2003/08/23/rss-scrape-urls",
  "prevPostTitle": "Scraping HTML with web services",
  "nextPostPath": "2003/08/28/css-rollovers",
  "nextPostTitle": "CSS, Background Images, and Rollovers"
}