ExtractThinker

ExtractThinker 项目介绍

ExtractThinker 是一个旨在通过大型语言模型（LLMs）从各种文件和文档中提取数据的库。它提供一种面向对象关系映射（ORM）风格的方式，让用户可以灵活高效地进行文档提取和处理工作流。

功能特色

多种文档加载支持：支持使用 Tesseract OCR、Azure Form Recognizer、AWS TextExtract 以及 Google Document AI 进行文档加载。
可定制提取：通过合同定义可以自定义提取内容。
异步处理：能够高效地处理大量文档。
多种文档格式支持：内置对多种文档格式的支持。
LLMs 与文件的 ORM 交互：实现 LLMs 与文件之间的面向对象关系映射交互。

安装方法

要安装 ExtractThinker，可以使用 pip 进行安装：

pip install extract_thinker

使用示例

下面是一个快速入门示例，展示了如何使用 Tesseract OCR 加载文档并提取合同中定义的特定字段。

import os
from dotenv import load_dotenv
from extract_thinker import DocumentLoaderTesseract, Extractor, Contract

load_dotenv()
cwd = os.getcwd()

class InvoiceContract(Contract):
    invoice_number: str
    invoice_date: str

tesseract_path = os.getenv("TESSERACT_PATH")
test_file_path = os.path.join(cwd, "test_images", "invoice.png")

extractor = Extractor()
extractor.load_document_loader(
    DocumentLoaderTesseract(tesseract_path)
)
extractor.load_llm("claude-3-haiku-20240307")

result = extractor.extract(test_file_path, InvoiceContract)

print("Invoice Number: ", result.invoice_number)
print("Invoice Date: ", result.invoice_date)

文件拆分示例

您还可以使用 ExtractThinker 分割和处理文档。以下是相关的实施方法：

import os
from dotenv import load_dotenv
from extract_thinker import DocumentLoaderTesseract, Extractor, Process, Classification, ImageSplitter

load_dotenv()

class DriverLicense(Contract):
    # Define your DriverLicense contract fields here
    pass

class InvoiceContract(Contract):
    invoice_number: str
    invoice_date: str

extractor = Extractor()
extractor.load_document_loader(DocumentLoaderTesseract(os.getenv("TESSERACT_PATH")))
extractor.load_llm("gpt-3.5-turbo")

classifications = [
    Classification(name="Driver License", description="This is a driver license", contract=DriverLicense, extractor=extractor),
    Classification(name="Invoice", description="This is an invoice", contract=InvoiceContract, extractor=extractor)
]

process = Process()
process.load_document_loader(DocumentLoaderTesseract(os.getenv("TESSERACT_PATH")))
process.load_splitter(ImageSplitter())

path = "..."

split_content = process.load_file(path)\
    .split(classifications)\
    .extract()

# Process the split_content as needed