← Docs · Payloads · Markdown · Agent register · Home

Docs

Document extraction for agents: download a public URL or pass base64, then return text/tables. PDF uses pdfplumber; empty/scanned pages fall back to poppler + Tesseract OCR with language packs (default hun+eng).

Service id: docs
Version: 0.1.0
Status: available

Canonical payloads: https://mcp.glc-rag.hu/guide/docs/payload · markdown · docs://docs/payload

Authentication

MCP endpoint: https://mcp.glc-rag.hu/mcp (streamable HTTP)

Agents (recommended): self-register with account_type=agent to get an auto-approved token — see https://mcp.glc-rag.hu/guide/agent.

Or register as a human on the public site (all listed services are auto-approved), wait for system-admin approval, then create a token.

Authorization: Bearer mcp_...

Cursor mcp.json example:

{
  "mcpServers": {
    "docs": {
      "url": "https://mcp.glc-rag.hu/mcp",
      "headers": {
        "Authorization": "Bearer mcp_YOUR_TOKEN"
      }
    }
  }
}

Tools

docs_extract

Extract text and tables from a document (pdf, docx, xlsx, txt, csv). Provide either a public http(s) URL or base64 content. For scanned PDFs, OCR runs automatically when little text is found.

Input schema:

{
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "description": "Public http(s) URL of the document"
    },
    "content_base64": {
      "type": "string",
      "description": "Raw document bytes as base64 (or data: URL)"
    },
    "filename": {
      "type": "string",
      "description": "Optional filename hint for type detection (e.g. report.pdf)"
    },
    "lang": {
      "type": "string",
      "description": "OCR languages, e.g. hun+eng (default hun+eng)",
      "default": "hun+eng"
    },
    "force_ocr": {
      "type": "boolean",
      "description": "Always OCR PDF pages even if embedded text exists",
      "default": false
    },
    "max_length": {
      "type": "integer",
      "minimum": 1,
      "maximum": 1000000,
      "default": 20000
    },
    "start_index": {
      "type": "integer",
      "minimum": 0,
      "default": 0
    },
    "max_pages": {
      "type": "integer",
      "minimum": 1,
      "maximum": 200,
      "default": 50,
      "description": "Max PDF pages to process"
    }
  },
  "additionalProperties": false
}

Examples:

{
  "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
}
{
  "url": "https://example.com/scan.pdf",
  "lang": "hun+eng",
  "force_ocr": true
}

docs_status

Docs service health, supported formats, and installed OCR languages.

Input schema:

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}

Examples:

{}

Usage notes

Call docs_extract with url or content_base64. Set lang to hun, eng, deu, or combinations like hun+eng. Use force_ocr: true for scanned PDFs. Paginate with max_length / start_index.

Errors / limits

Unsupported type, SSRF-blocked URL, oversized body, missing OCR packs, or extract failure → {error, is_error}. Invalid API key → platform 401.

Agent discovery