alpcentaur
|
d3335f203b
|
added trafilatura exception
|
2023-11-22 00:02:29 +00:00 |
|
alpcentaur
|
14ece9bceb
|
added functions for uniform and not uniform entry end points - non uniform endpoints are generally parsed as text from any paragraph xml element p
|
2023-11-20 15:28:04 +00:00 |
|
alpcentaur
|
42841ee650
|
added some exceptions for bad encoding and get errors
|
2023-11-14 14:38:45 +00:00 |
|
alpcentaur
|
317ef99720
|
changed code in entrylist data2dictionary to handle empty or missing xml elements
|
2023-11-14 10:22:26 +00:00 |
|
alpcentaur
|
ff23c22e3c
|
added working bund.de-bekanntmachungen config with new example of xpath contains
|
2023-11-13 16:44:11 +00:00 |
|
alpcentaur
|
06fa81e549
|
added function find config parameter and changed core spider
|
2023-11-10 01:12:49 +00:00 |
|
alpcentaur
|
a846ce04cc
|
specifying the links, new exception clause if soupparser does not work
|
2023-11-07 14:55:05 +00:00 |
|
alpcentaur
|
c078ee4b1b
|
first function works, actuall xml parser has still problems with certain xml types
|
2023-11-06 19:17:45 +00:00 |
|