> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.athenaintel.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.athenaintel.com/_mcp/server.

# Structured Output

> Extract structured data from text using JSON schemas.

Use the Structured Data Extractor to extract structured data from text chunks using JSON schemas. This is useful for parsing unstructured text into well-defined data structures.

Key features:

* Define output structure using JSON Schema (draft 2020-12)
* Process multiple text chunks with map-reduce pattern
* Get validated, structured output matching your schema

### Install Package

#### Jupyter Notebook

```python
!pip install -U athenaintel
```

#### Terminal/Shell

```bash
pip install -U athenaintel
```

### Set Up Client

```python
from athena import Athena, Chunk, ChunkContentItem_Text

athena = Athena(api_key="<YOUR_API_KEY>")
```

### Define Your Schema

Create a JSON schema that describes the structure you want to extract:

```python
person_schema = {
    "title": "Person",
    "description": "Information about a person",
    "type": "object",
    "properties": {
        "name": {"type": "string"},
        "age": {"type": "integer"},
        "email": {"type": "string"}
    },
    "required": ["name"]
}
```

### Extract Structured Data

Pass text chunks and your schema to the structured data extractor:

```python
response = athena.tools.structured_data_extractor.invoke(
    chunks=[
        Chunk(
            chunk_id="1",
            content=[
                ChunkContentItem_Text(
                    text="John Smith is a 35 year old software developer. Contact him at john.smith@example.com"
                )
            ]
        ),
        Chunk(
            chunk_id="2",
            content=[
                ChunkContentItem_Text(
                    text="Jane Doe is a 28 year old data scientist. Her email is jane.doe@example.com"
                )
            ]
        )
    ],
    json_schema=person_schema,
    reduce=True
)

print(response.reduced_data)
```

```json
{
    "name": "John Smith",
    "age": 35,
    "email": "john.smith@example.com"
}
```

### Access Chunk-by-Chunk Results

To get extracted data from each chunk individually, set `reduce=False`:

```python
response = athena.tools.structured_data_extractor.invoke(
    chunks=[
        Chunk(
            chunk_id="1",
            content=[
                ChunkContentItem_Text(
                    text="John Smith is a 35 year old software developer."
                )
            ]
        ),
        Chunk(
            chunk_id="2",
            content=[
                ChunkContentItem_Text(
                    text="Jane Doe is a 28 year old data scientist."
                )
            ]
        )
    ],
    json_schema=person_schema,
    reduce=False
)

for chunk_result in response.chunk_by_chunk_data:
    print(f"Chunk {chunk_result.chunk_id}: {chunk_result.data}")
```

```
Chunk 1: {'name': 'John Smith', 'age': 35}
Chunk 2: {'name': 'Jane Doe', 'age': 28}
```