This role sits close to the part of VerifyPDF that matters most: deciding whether a submitted document can be trusted. We already check metadata, PDF structure, fonts, incremental updates, suspicious producers, OCR output, visual signals and template matches. We want someone who can make those signals better, test them properly and turn messy evidence into a verdict customers can understand.
Key Responsibilities:
You'd work on document classification, signal extraction, evaluation sets, prompt and model behavior, false-positive review and the scoring path that produces trust scores and risk bands. Some weeks this is applied ML. Some weeks it is reading strange PDFs, finding why our pipeline missed something and writing the fix yourself.
You'd build tools that make fraud analysis less subjective: repeatable evals, regression tests for known fraud patterns, better evidence summaries and safer fallbacks when OCR or model output is weak. You'd work with Python, FastAPI, pdfplumber, pdfminer, PyMuPDF, OpenAI, SageMaker inference, scikit-learn and the existing PDF forensics code.
You'd also help decide when AI should not be used. If a deterministic PDF check is more reliable than a model, we should use the deterministic check. The goal is not to add AI everywhere. The goal is to catch more fraud and explain the result better.