Analyze the collected literature: extract structured information, embed and cluster paper abstracts, and classify papers with local LLMs. The classification scripts also tune prompts and produce the labeled modeling_papers.json outputs.
Part of the genscai use-case series. Shared library code lives in the genscai package; datasets in data/ and generated artifacts in output/ are shared at the repo root.