72 Commits (master)
 

Author SHA1 Message Date
  alpcentaur df4a8289b8 added pdf parser if entry link is direct pdf 9 months ago
  alpcentaur 677e54c0c2 added trafilatura to requirements 9 months ago
  alpcentaur 9ceaa28a82 Merge remote-tracking branch 'refs/remotes/origin/master' 9 months ago
  alpcentaur d3335f203b added trafilatura exception 9 months ago
  alpcentaur 61f9ba67fb update README.md 10 months ago
  alpcentaur 69c517292b Update 'README.md' 10 months ago
  alpcentaur 14ece9bceb added functions for uniform and not uniform entry end points - non uniform endpoints are generally parsed as text from any paragraph xml element p 10 months ago
  alpcentaur b2cf4b67ce added first config parameters for search on not uniform entries 10 months ago
  alpcentaur 42841ee650 added some exceptions for bad encoding and get errors 10 months ago
  alpcentaur 317ef99720 changed code in entrylist data2dictionary to handle empty or missing xml elements 10 months ago
  alpcentaur ff23c22e3c added working bund.de-bekanntmachungen config with new example of xpath contains 10 months ago
  alpcentaur 06fa81e549 added function find config parameter and changed core spider 10 months ago
  alpcentaur a846ce04cc specifying the links, new exception clause if soupparser does not work 10 months ago
  alpcentaur a99881796a first function works, actuall xml parser has still problems with certain xml types 10 months ago
  alpcentaur c078ee4b1b first function works, actuall xml parser has still problems with certain xml types 10 months ago
  alpcentaur 8b20bc178f added multi pages configuration and code 10 months ago
  alpcentaur 7aa903883b update to config.yaml 10 months ago
  alpcentaur 59838bb8e1 added main.py importing and using the spider functions 10 months ago
  alpcentaur 5ac07d151a added first config.yaml template and started creating folder structure 10 months ago
  alpcentaur b3011efc73 small change of naming in error message added 10 months ago
  alpcentaur 687d40f156 first change of naming, first commit for the actual spider based on importPEP 10 months ago
  alpcentaur 8783251133 first commit 10 months ago