Tried fine-tuning a tiny model for my shop's invoices, got a pleasant surprise
I spent about 3 days feeding my old invoice data into a small open-source model to auto-categorize parts and labor, mostly as a weekend experiment. It nailed 94% of the line items on the first run, and I only had to fix about 20 entries out of 330. Has anyone else found that smaller, focused models beat the big ones for niche tasks like this?
Man, 94% on the first run is wild, and it makes total sense. The big models have to be good at everything, so they spread themselves thin. A tiny model that only ever sees invoice lines from one shop starts to pick up on your specific patterns, like the way you always write "1/2 in copper elbow" or how your labor lines always start with a verb. Keep going with it and you could push that 94 up by feeding it the 20 you had to fix, since those corrections are basically the model learning your edge cases. The small model also runs on cheap hardware and spits out answers fast, which matters more than raw smarts when you are just sorting line items. That focused setup is probably why it beat the big ones for your task.