UsageΒΆ
The package exposes a single high-level entry point for end-to-end parsing:
from sci_fi_parser import parse_folder
result = parse_folder(
"path/to/pdfs",
output_dir="output",
extracted_image_dir="temp/extracted_images",
vlm_config="config/vlm.toml",
)
print(result.summary())
parse_folder() walks a PDF file or a flat directory of PDFs, extracts chart
images, classifies them, runs OCR / computer vision, and finally applies the
VLM stage by default. If output_dir is provided, the parsed dataset is written to JSONL
and Parquet outputs. Alternatively, the output can be saved using the returned object with ParseResult.save()
The returned ParseResult gives access to the parsed
image and PDF collections:
image_set = result.images
pdf_set = result.pdfs
for image_id, record in image_set.items():
print(image_id, record)
Saved output is organized as:
raw/image_set.jsonlfor the full record streamtables/charts.parquetfor one row per charttables/series.parquetfor one row per seriestables/points.parquetfor one row per extracted point
For batch runs, keep the extraction folder separate from the final output directory. The extraction folder is temporary working storage and is not meant to be checked in or treated as a final artifact.