Skip to navigation

Structured Output

Use the Structured Data Extractor to extract structured data from text chunks using JSON schemas. This is useful for parsing unstructured text into well-defined data structures.

Key features:

  • Define output structure using JSON Schema (draft 2020-12)
  • Process multiple text chunks with map-reduce pattern
  • Get validated, structured output matching your schema
1

Install Package

!pip install -U athenaintel
2

Set Up Client

from athena import Athena, Chunk, ChunkContentItem_Text
athena = Athena(api_key="<YOUR_API_KEY>")
3

Define Your Schema

Create a JSON schema that describes the structure you want to extract:

person_schema = {
"title": "Person",
"description": "Information about a person",
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
"email": {"type": "string"}
},
"required": ["name"]
}
4

Extract Structured Data

Pass text chunks and your schema to the structured data extractor:

response = athena.tools.structured_data_extractor.invoke(
chunks=[
Chunk(
chunk_id="1",
content=[
ChunkContentItem_Text(
text="John Smith is a 35 year old software developer. Contact him at john.smith@example.com"
)
]
),
Chunk(
chunk_id="2",
content=[
ChunkContentItem_Text(
text="Jane Doe is a 28 year old data scientist. Her email is jane.doe@example.com"
)
]
)
],
json_schema=person_schema,
reduce=True
)
print(response.reduced_data)
{
"name": "John Smith",
"age": 35,
"email": "john.smith@example.com"
}
5

Access Chunk-by-Chunk Results

To get extracted data from each chunk individually, set reduce=False:

response = athena.tools.structured_data_extractor.invoke(
chunks=[
Chunk(
chunk_id="1",
content=[
ChunkContentItem_Text(
text="John Smith is a 35 year old software developer."
)
]
),
Chunk(
chunk_id="2",
content=[
ChunkContentItem_Text(
text="Jane Doe is a 28 year old data scientist."
)
]
)
],
json_schema=person_schema,
reduce=False
)
for chunk_result in response.chunk_by_chunk_data:
print(f"Chunk {chunk_result.chunk_id}: {chunk_result.data}")
Chunk 1: {'name': 'John Smith', 'age': 35}
Chunk 2: {'name': 'Jane Doe', 'age': 28}