> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.athenaintel.com/python-guides/structured-output/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.athenaintel.com/_mcp/server. # Structured Output > Extract structured data from text using JSON schemas. Use the Structured Data Extractor to extract structured data from text chunks using JSON schemas. This is useful for parsing unstructured text into well-defined data structures. Key features: * Define output structure using JSON Schema (draft 2020-12) * Process multiple text chunks with map-reduce pattern * Get validated, structured output matching your schema ### Install Package #### Jupyter Notebook ```python !pip install -U athenaintel ``` #### Terminal/Shell ```bash pip install -U athenaintel ``` ### Set Up Client ```python from athena import Athena, Chunk, ChunkContentItem_Text athena = Athena(api_key="") ``` ### Define Your Schema Create a JSON schema that describes the structure you want to extract: ```python person_schema = { "title": "Person", "description": "Information about a person", "type": "object", "properties": { "name": {"type": "string"}, "age": {"type": "integer"}, "email": {"type": "string"} }, "required": ["name"] } ``` ### Extract Structured Data Pass text chunks and your schema to the structured data extractor: ```python response = athena.tools.structured_data_extractor.invoke( chunks=[ Chunk( chunk_id="1", content=[ ChunkContentItem_Text( text="John Smith is a 35 year old software developer. Contact him at john.smith@example.com" ) ] ), Chunk( chunk_id="2", content=[ ChunkContentItem_Text( text="Jane Doe is a 28 year old data scientist. Her email is jane.doe@example.com" ) ] ) ], json_schema=person_schema, reduce=True ) print(response.reduced_data) ``` ```json { "name": "John Smith", "age": 35, "email": "john.smith@example.com" } ``` ### Access Chunk-by-Chunk Results To get extracted data from each chunk individually, set `reduce=False`: ```python response = athena.tools.structured_data_extractor.invoke( chunks=[ Chunk( chunk_id="1", content=[ ChunkContentItem_Text( text="John Smith is a 35 year old software developer." ) ] ), Chunk( chunk_id="2", content=[ ChunkContentItem_Text( text="Jane Doe is a 28 year old data scientist." ) ] ) ], json_schema=person_schema, reduce=False ) for chunk_result in response.chunk_by_chunk_data: print(f"Chunk {chunk_result.chunk_id}: {chunk_result.data}") ``` ``` Chunk 1: {'name': 'John Smith', 'age': 35} Chunk 2: {'name': 'Jane Doe', 'age': 28} ``` > Extract structured data from text using JSON schemas.