Skip to content

MD scraping problems #785

Description

@stucka

MD assumes there's a single header for more than a year of data. The state recently changed the format of the HTML, and the assumption that none of the organization will change seems unsafe from a sanitary perspective.

Ideally, each HTML file should be parsed independently such that each row of data is matched with headers from that file. This may require some standardization/lookups but can make things go much easier in the future.

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions