Schematron-3B v1.0
Open Source
Infrastructure
Schematron-3B is a long-context extraction model designed to convert noisy HTML into clean, typed JSON that adheres to a user-provided schema. It is specifically trained for web scraping, data ingestion, and transforming arbitrary web pages into structured records.
Schema-first extraction: Outputs strictly conforming JSON to the provided schema
Long context: Handles lengthy, noisy HTML up to 128K tokens
Variants: Available in 3B (default, cost-efficient) and 8B (higher quality) sizes