A supplier sends you a PDF. You need article numbers, descriptions and prices in your software. The PDF has merged cells, prices per 100 pieces, and a footnote explaining that the number you just extracted means something else entirely.
This is the kind of AI application i find interesting. Small problem, concrete output, very little room for creative writing.
My price-list demo on Hugging Face turns PDFs and images into JSON, CSV and DATANORM 4. It uses stock Qwen3-VL-8B-Instruct with document preprocessing and validation. That working extraction pipeline is the starting point here: what it does, what a fine-tuning dataset needs to contain, and how to train a LoRA adapter against the same task.
From PDF to JSON
The model gets a page image or a smaller table section, together with the extraction instruction. Dense tables with clear horizontal rules are split into sections while headers and footnotes remain available as context.
That context matters. A price per 100 pieces and a box containing 100 pieces are different fields. Mix those up and your perfectly valid JSON is still wrong.
1 | PDF page or table section |
Processing produces checkpoints so a quota limit doesn’t mean starting the document again. Resuming requires the same model revision and extraction instruction. Otherwise one result could quietly contain answers from two different extraction setups.
Where a PDF has a usable text layer, ordinary code also checks explicit VAT wording and narrowly recognizable packaging columns. Conflicts stay visible for review. The original price text is retained alongside the normalized value, including after reloading a checkpoint.
The model proposes data. Pydantic checks the structure. The exporter checks what it can represent and writes the file without another GPU call. Better cropping and better checks can improve that application without changing a single model weight.
Decide what the model should learn
The model returns price-rows-v2: eleven fixed fields per row, expanded into the full JSON contract by Python. Repeating long property names for every article uses output tokens without adding information.
For a fictional row reading AB-1042 | Sechskantschraube M8 | 12,50 EUR / 100 Stück, a compact target could look like this:
1 | { |
The fields are article number, description, variant, price, currency, price quantity, unit, tax basis, pack quantity, minimum quantity and source excerpt. Tax stays unknown because that fictional row doesn’t establish it. The 100 belongs to the price quantity, not the packaging.
The current transport allows at most 120 rows per response. Hitting that limit means the input needs splitting; it doesn’t make the remaining articles disappear successfully.
Use this same compact format for the training targets if this is what inference expects. Then validate the expanded output. Teaching verbose objects and deploying positional arrays are two different tasks, even if both happen to be JSON.
Preparing the fine-tuning dataset
A training example pairs the page image or section, its extraction instruction, and the corrected answer. Supervised fine-tuning trains the model to produce that answer from the input.
Include awkward examples: unreadable cells, repeated headers, variants, quantity breaks and pages with no articles. Missing prices stay null. Training only on clean tables teaches the model that every page has a convenient answer.
Model-generated labels can give you a head start. They still need checking against the source. Feeding uncorrected output back into training is a good way to make the same mistake more confidently.
Split by supplier or document family before training. Keep sections and pages from the same catalogue together. A random crop split can put almost identical examples into training and evaluation, then reward you for memorizing them.
Run the stock model on the held-out documents too, using the same preprocessing and validation as the adapter. Otherwise you won’t know whether an improvement came from training or from fixing the table cropping.
Fine-tuning Qwen3-VL with LoRA
You don’t need to pretrain a vision model from scratch. LoRA adds trainable low-rank adapters while keeping the original weights frozen. You still load the base model; the adapter contains the learned changes.
Hugging Face’s TRL SFTTrainer supports image datasets and PEFT adapters. The training core can look like this:
1 | from peft import LoraConfig |
This is an illustrative training configuration to build on, rather than a measured recipe for this dataset. train_ds must contain decoded page images in an image column and conversational prompt/completion fields: the user prompt includes the image placeholder and extraction instruction; the assistant completion contains the corrected compact JSON shown above. TRL’s prompt-completion format trains on the completion by default.
Use a CUDA training environment with BF16 support and enough VRAM. Pin compatible versions of PyTorch, Transformers, TRL and PEFT, and record the base-model revision with your dataset revision. Install those training dependencies separately from the demo’s serving environment.
One detail worth paying attention to: blindly truncating a multimodal sample can cut out image tokens. max_length=None avoids that truncation, but doesn’t make memory unlimited. Bound image resolution and document size, then run a small batch through the whole training path before committing to a longer job. If memory is tight, investigate QLoRA with a quantized base model rather than assuming this example fits your GPU.
Evaluate the extracted prices, not just the loss
Run the base model and adapter on the same held-out documents. Match the extracted rows to the reference rows and compare article numbers, exact prices and their quantity basis. Count missing rows and invented rows too. A lower training loss doesn’t tell you whether 12.50 EUR / 100 Stück became 12.50 EUR / Stück.
Keep a validation set for choosing checkpoints and a separate test set for the final comparison. Only put the adapter into the application if it earns its place there. Deployment also needs the adapter loaded on top of the matching base model, or deliberately merged weights; changing a model-name string isn’t enough for the current Space loader.
And keep the DATANORM writer as normal code. A model can help read a messy table. It doesn’t need to improvise the file format your accounting software imports.
If you want to poke at it yourself, try the demo on Hugging Face. Pick one of the sample price lists and compare the JSON with the page. See where it holds up, and where it still needs a human.