{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# German Document OCR with Qwen2-VL\n",
    "\n",
    "This cookbook demonstrates how to use Qwen2-VL for Optical Character Recognition (OCR) on **German documents**. German business documents have unique characteristics that require special attention:\n",
    "\n",
    "- **Umlauts** (ä, ö, ü) and **Eszett** (ß)\n",
    "- **German date formats** (DD.MM.YYYY)\n",
    "- **Currency formatting** (1.234,56 €)\n",
    "- **Specific document types**: Rechnungen (invoices), Formulare (forms), Ausweise (IDs)\n",
    "\n",
    "## Use Cases Covered\n",
    "\n",
    "1. **Invoice OCR** (Rechnungserkennung)\n",
    "2. **Form Processing** (Formularverarbeitung)\n",
    "3. **ID Document Extraction** (Ausweiserkennung)\n",
    "4. **Structured Data Extraction** with JSON output\n",
    "\n",
    "---\n",
    "\n",
    "**Author:** [Keyvan Hardani](https://keyvan.ai) - AI Engineer specializing in German Document Intelligence\n",
    "\n",
    "**Related Project:** [German-OCR](https://github.com/Keyvanhardani/German-OCR) - Specialized OCR for German documents"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Setup\n",
    "\n",
    "First, install the required dependencies:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "!pip install transformers torch qwen-vl-utils accelerate -q"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import torch\n",
    "from transformers import Qwen2VLForConditionalGeneration, AutoProcessor\n",
    "from qwen_vl_utils import process_vision_info\n",
    "import json\n",
    "import re\n",
    "from PIL import Image\n",
    "import requests\n",
    "from io import BytesIO"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Load the Model\n",
    "\n",
    "We use `Qwen2-VL-7B-Instruct` which offers excellent multilingual OCR capabilities including German."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Load model and processor\n",
    "model_name = \"Qwen/Qwen2-VL-7B-Instruct\"\n",
    "\n",
    "model = Qwen2VLForConditionalGeneration.from_pretrained(\n",
    "    model_name,\n",
    "    torch_dtype=torch.bfloat16,\n",
    "    attn_implementation=\"flash_attention_2\",\n",
    "    device_map=\"auto\"\n",
    ")\n",
    "\n",
    "processor = AutoProcessor.from_pretrained(model_name)\n",
    "\n",
    "print(f\"Model loaded: {model_name}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Helper Functions"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "def load_image(image_path_or_url: str) -> Image.Image:\n",
    "    \"\"\"Load image from local path or URL.\"\"\"\n",
    "    if image_path_or_url.startswith(('http://', 'https://')):\n",
    "        response = requests.get(image_path_or_url)\n",
    "        return Image.open(BytesIO(response.content)).convert('RGB')\n",
    "    return Image.open(image_path_or_url).convert('RGB')\n",
    "\n",
    "\n",
    "def extract_json(text: str) -> dict:\n",
    "    \"\"\"Extract JSON from model response (handles markdown code blocks).\"\"\"\n",
    "    # Try to find JSON in markdown code block\n",
    "    json_match = re.search(r'```(?:json)?\\s*([\\s\\S]*?)```', text)\n",
    "    if json_match:\n",
    "        text = json_match.group(1)\n",
    "    \n",
    "    # Clean and parse\n",
    "    text = text.strip()\n",
    "    try:\n",
    "        return json.loads(text)\n",
    "    except json.JSONDecodeError:\n",
    "        return {\"raw_text\": text, \"parse_error\": True}\n",
    "\n",
    "\n",
    "def run_ocr(image_source: str, prompt: str, max_tokens: int = 2048) -> str:\n",
    "    \"\"\"Run OCR inference on an image with a custom prompt.\"\"\"\n",
    "    messages = [\n",
    "        {\n",
    "            \"role\": \"user\",\n",
    "            \"content\": [\n",
    "                {\"type\": \"image\", \"image\": image_source},\n",
    "                {\"type\": \"text\", \"text\": prompt}\n",
    "            ]\n",
    "        }\n",
    "    ]\n",
    "    \n",
    "    # Prepare inputs\n",
    "    text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)\n",
    "    image_inputs, video_inputs = process_vision_info(messages)\n",
    "    \n",
    "    inputs = processor(\n",
    "        text=[text],\n",
    "        images=image_inputs,\n",
    "        videos=video_inputs,\n",
    "        padding=True,\n",
    "        return_tensors=\"pt\"\n",
    "    ).to(model.device)\n",
    "    \n",
    "    # Generate\n",
    "    with torch.no_grad():\n",
    "        generated_ids = model.generate(\n",
    "            **inputs,\n",
    "            max_new_tokens=max_tokens,\n",
    "            do_sample=False\n",
    "        )\n",
    "    \n",
    "    # Decode\n",
    "    generated_ids_trimmed = [\n",
    "        out_ids[len(in_ids):] \n",
    "        for in_ids, out_ids in zip(inputs.input_ids, generated_ids)\n",
    "    ]\n",
    "    \n",
    "    return processor.batch_decode(\n",
    "        generated_ids_trimmed, \n",
    "        skip_special_tokens=True, \n",
    "        clean_up_tokenization_spaces=False\n",
    "    )[0]"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## 1. German Invoice OCR (Rechnungserkennung)\n",
    "\n",
    "German invoices (\"Rechnungen\") contain specific fields that are legally required:\n",
    "\n",
    "- **Rechnungsnummer** (Invoice number)\n",
    "- **Rechnungsdatum** (Invoice date)\n",
    "- **Steuernummer / USt-IdNr.** (Tax ID / VAT number)\n",
    "- **Nettobetrag, MwSt., Bruttobetrag** (Net, VAT, Gross amounts)\n",
    "- **IBAN / BIC** (Bank details)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# German Invoice Extraction Prompt\n",
    "GERMAN_INVOICE_PROMPT = \"\"\"\n",
    "Analysiere diese deutsche Rechnung und extrahiere alle relevanten Informationen.\n",
    "\n",
    "Gib die Daten als JSON mit folgender Struktur zurück:\n",
    "\n",
    "```json\n",
    "{\n",
    "    \"rechnungsnummer\": \"string\",\n",
    "    \"rechnungsdatum\": \"DD.MM.YYYY\",\n",
    "    \"lieferdatum\": \"DD.MM.YYYY oder null\",\n",
    "    \"absender\": {\n",
    "        \"firma\": \"string\",\n",
    "        \"adresse\": \"string\",\n",
    "        \"steuernummer\": \"string oder null\",\n",
    "        \"ust_idnr\": \"string oder null\"\n",
    "    },\n",
    "    \"empfaenger\": {\n",
    "        \"name\": \"string\",\n",
    "        \"adresse\": \"string\"\n",
    "    },\n",
    "    \"positionen\": [\n",
    "        {\n",
    "            \"beschreibung\": \"string\",\n",
    "            \"menge\": number,\n",
    "            \"einzelpreis\": number,\n",
    "            \"gesamtpreis\": number\n",
    "        }\n",
    "    ],\n",
    "    \"betraege\": {\n",
    "        \"netto\": number,\n",
    "        \"mwst_satz\": number,\n",
    "        \"mwst_betrag\": number,\n",
    "        \"brutto\": number\n",
    "    },\n",
    "    \"zahlungsinformationen\": {\n",
    "        \"iban\": \"string oder null\",\n",
    "        \"bic\": \"string oder null\",\n",
    "        \"zahlungsziel\": \"string oder null\"\n",
    "    }\n",
    "}\n",
    "```\n",
    "\n",
    "Wichtig:\n",
    "- Behalte das deutsche Datumsformat (DD.MM.YYYY)\n",
    "- Konvertiere Beträge zu Zahlen (1.234,56 € → 1234.56)\n",
    "- Setze fehlende Felder auf null\n",
    "\"\"\"\n",
    "\n",
    "# Example usage (replace with your invoice image)\n",
    "# invoice_image = \"path/to/german_invoice.jpg\"\n",
    "# result = run_ocr(invoice_image, GERMAN_INVOICE_PROMPT)\n",
    "# invoice_data = extract_json(result)\n",
    "# print(json.dumps(invoice_data, indent=2, ensure_ascii=False))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Example: Process a Sample Invoice"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Demo with a sample German invoice\n",
    "# You can replace this URL with your own invoice image\n",
    "\n",
    "sample_invoice_url = \"YOUR_INVOICE_IMAGE_URL_HERE\"\n",
    "\n",
    "# Uncomment to run:\n",
    "# result = run_ocr(sample_invoice_url, GERMAN_INVOICE_PROMPT)\n",
    "# invoice_data = extract_json(result)\n",
    "# \n",
    "# print(\"=\" * 50)\n",
    "# print(\"EXTRAHIERTE RECHNUNGSDATEN\")\n",
    "# print(\"=\" * 50)\n",
    "# print(json.dumps(invoice_data, indent=2, ensure_ascii=False))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## 2. German Form Processing (Formularverarbeitung)\n",
    "\n",
    "German forms often include:\n",
    "- Checkboxes (Ankreuzfelder)\n",
    "- Handwritten entries\n",
    "- Structured fields with labels"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# German Form Extraction Prompt\n",
    "GERMAN_FORM_PROMPT = \"\"\"\n",
    "Analysiere dieses deutsche Formular und extrahiere alle ausgefüllten Felder.\n",
    "\n",
    "Gib die Daten als JSON zurück mit:\n",
    "- Feldname als Schlüssel\n",
    "- Eingetragener Wert als Wert\n",
    "- Bei Ankreuzfeldern: true/false\n",
    "- Bei leeren Feldern: null\n",
    "\n",
    "Beispiel:\n",
    "```json\n",
    "{\n",
    "    \"vorname\": \"Max\",\n",
    "    \"nachname\": \"Mustermann\",\n",
    "    \"geburtsdatum\": \"15.03.1985\",\n",
    "    \"geschlecht_maennlich\": true,\n",
    "    \"geschlecht_weiblich\": false,\n",
    "    \"telefon\": null\n",
    "}\n",
    "```\n",
    "\n",
    "Erkenne auch handschriftliche Einträge.\n",
    "\"\"\"\n",
    "\n",
    "# Example usage:\n",
    "# form_image = \"path/to/german_form.jpg\"\n",
    "# result = run_ocr(form_image, GERMAN_FORM_PROMPT)\n",
    "# form_data = extract_json(result)\n",
    "# print(json.dumps(form_data, indent=2, ensure_ascii=False))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## 3. ID Document Extraction (Ausweiserkennung)\n",
    "\n",
    "Extract information from German ID cards (Personalausweis) and passports.\n",
    "\n",
    "**Note:** Always handle personal data according to GDPR (DSGVO) regulations!"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# German ID Extraction Prompt\n",
    "GERMAN_ID_PROMPT = \"\"\"\n",
    "Extrahiere die Informationen aus diesem deutschen Ausweisdokument.\n",
    "\n",
    "Gib die Daten als JSON zurück:\n",
    "\n",
    "```json\n",
    "{\n",
    "    \"dokumenttyp\": \"Personalausweis/Reisepass/Führerschein\",\n",
    "    \"nachname\": \"string\",\n",
    "    \"vorname\": \"string\",\n",
    "    \"geburtsdatum\": \"DD.MM.YYYY\",\n",
    "    \"geburtsort\": \"string\",\n",
    "    \"nationalitaet\": \"string\",\n",
    "    \"ausweisnummer\": \"string\",\n",
    "    \"gueltig_bis\": \"DD.MM.YYYY\",\n",
    "    \"ausstellende_behoerde\": \"string oder null\"\n",
    "}\n",
    "```\n",
    "\n",
    "Hinweis: Achte auf korrekte Umlaute (ä, ö, ü, ß).\n",
    "\"\"\"\n",
    "\n",
    "# Example usage:\n",
    "# id_image = \"path/to/german_id.jpg\"\n",
    "# result = run_ocr(id_image, GERMAN_ID_PROMPT)\n",
    "# id_data = extract_json(result)\n",
    "# print(json.dumps(id_data, indent=2, ensure_ascii=False))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## 4. Full-Page German OCR\n",
    "\n",
    "For general German text extraction without structured output:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Simple German OCR Prompt\n",
    "GERMAN_OCR_SIMPLE = \"\"\"\n",
    "Extrahiere den gesamten Text aus diesem Bild.\n",
    "\n",
    "Regeln:\n",
    "- Behalte die ursprüngliche Formatierung bei (Absätze, Listen)\n",
    "- Achte auf korrekte deutsche Zeichen (ä, ö, ü, ß)\n",
    "- Erkenne Tabellen und formatiere sie lesbar\n",
    "- Gib nur den extrahierten Text zurück, keine Erklärungen\n",
    "\"\"\"\n",
    "\n",
    "# Example usage:\n",
    "# document_image = \"path/to/german_document.jpg\"\n",
    "# extracted_text = run_ocr(document_image, GERMAN_OCR_SIMPLE)\n",
    "# print(extracted_text)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## 5. Batch Processing Multiple Documents"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "def process_german_documents(image_paths: list, document_type: str = \"invoice\") -> list:\n",
    "    \"\"\"\n",
    "    Process multiple German documents in batch.\n",
    "    \n",
    "    Args:\n",
    "        image_paths: List of image file paths or URLs\n",
    "        document_type: One of 'invoice', 'form', 'id', 'general'\n",
    "    \n",
    "    Returns:\n",
    "        List of extracted data dictionaries\n",
    "    \"\"\"\n",
    "    prompts = {\n",
    "        \"invoice\": GERMAN_INVOICE_PROMPT,\n",
    "        \"form\": GERMAN_FORM_PROMPT,\n",
    "        \"id\": GERMAN_ID_PROMPT,\n",
    "        \"general\": GERMAN_OCR_SIMPLE\n",
    "    }\n",
    "    \n",
    "    prompt = prompts.get(document_type, GERMAN_OCR_SIMPLE)\n",
    "    results = []\n",
    "    \n",
    "    for i, image_path in enumerate(image_paths):\n",
    "        print(f\"Processing document {i+1}/{len(image_paths)}: {image_path}\")\n",
    "        try:\n",
    "            result = run_ocr(image_path, prompt)\n",
    "            if document_type != \"general\":\n",
    "                result = extract_json(result)\n",
    "            results.append({\"file\": image_path, \"data\": result, \"success\": True})\n",
    "        except Exception as e:\n",
    "            results.append({\"file\": image_path, \"error\": str(e), \"success\": False})\n",
    "    \n",
    "    return results\n",
    "\n",
    "# Example usage:\n",
    "# documents = [\"invoice1.jpg\", \"invoice2.jpg\", \"invoice3.jpg\"]\n",
    "# results = process_german_documents(documents, document_type=\"invoice\")\n",
    "# for r in results:\n",
    "#     print(json.dumps(r, indent=2, ensure_ascii=False))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Best Practices for German Document OCR\n",
    "\n",
    "### 1. Image Quality\n",
    "- **Resolution**: Minimum 150 DPI, ideally 300 DPI\n",
    "- **Lighting**: Even lighting, avoid shadows\n",
    "- **Orientation**: Correct rotation before processing\n",
    "\n",
    "### 2. German-Specific Considerations\n",
    "- Always validate Umlauts (ä, ö, ü) and Eszett (ß)\n",
    "- German dates use DD.MM.YYYY format\n",
    "- Currency: Decimal comma (1.234,56 €)\n",
    "- Tax IDs follow specific patterns (DE + 9 digits for VAT)\n",
    "\n",
    "### 3. GDPR/DSGVO Compliance\n",
    "- Process personal data locally when possible\n",
    "- Implement data minimization\n",
    "- Log access to sensitive documents\n",
    "- Delete processed images after extraction\n",
    "\n",
    "### 4. Error Handling\n",
    "- Validate extracted dates and amounts\n",
    "- Implement confidence scores\n",
    "- Manual review for low-confidence extractions"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Resources\n",
    "\n",
    "- **German-OCR Project**: [github.com/Keyvanhardani/German-OCR](https://github.com/Keyvanhardani/German-OCR) - Specialized OCR fine-tuned for German documents\n",
    "- **Qwen2-VL Documentation**: [Hugging Face Model Card](https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct)\n",
    "- **DSGVO Guidelines**: [GDPR compliance for document processing](https://gdpr.eu/)\n",
    "\n",
    "---\n",
    "\n",
    "*This cookbook was created by [Keyvan Hardani](https://keyvan.ai) to help the German-speaking community leverage Qwen2-VL for document processing.*"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python",
   "version": "3.10.0"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}
