> ## Documentation Index
> Fetch the complete documentation index at: https://docs.macroscope.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Export schema

> S3 object layout, window manifests, and NDJSON record fields.

[S3 log export](/s3-log-export) writes gzip-compressed NDJSON parts and a manifest for each time window. Read the manifest first and consume only the parts it lists.

## File layout

Every export writes the same batched, gzip-compressed **NDJSON-per-window** layout, uniform across all three log types. Three concepts are enough to consume any of them:

* **Windows**: each scheduled run exports one time window per enabled log type.
* **Part files:** The window's records are written as gzip-compressed NDJSON (one JSON object per line), rolled into one or more part files of up to \~128 MB compressed each.
* **Manifest**: a single `.manifest.json` per window lists the authoritative set of part files. It is written last, as the atomic commit point.

### Object key layout

All three log types share the same key structure. Each window emits part files and a sibling manifest:

```text theme={null}
{key_prefix}{workspace_type}/{workspace_id}/{type}/{YYYY}/{MM}/{DD}/{type}_{window_start}_{window_end}.part-NNNN.ndjson.gz
{key_prefix}{workspace_type}/{workspace_id}/{type}/{YYYY}/{MM}/{DD}/{type}_{window_start}_{window_end}.manifest.json
```

| Segment | Description |
| - | - |
| `{key_prefix}` | Your optional key prefix (empty if unset; ends in `/` if set). |
| `{workspace_type}/{workspace_id}` | Scopes every object to your workspace, so multiple workspaces can safely share one bucket and prefix. |
| `{type}` | One of `check-run-logs`, `code-review`, or `activity-logs`. |
| `{YYYY}/{MM}/{DD}` | Date partition, derived from the window's start time (UTC). |
| `{type}_{window_start}_{window_end}` | Filename stem: the log type token, then the window bounds in compact ISO 8601 UTC, e.g. `check-run-logs_20260716T180000Z_20260716T190000Z`. |
| `.part-NNNN.ndjson.gz` | A gzipped-NDJSON part file. `NNNN` is a 1-based, zero-padded ordinal (`0001`, `0002`, ...). |
| `.manifest.json` | The window manifest (see below). |

Records roll into a new part file once the current part reaches roughly 128 MB compressed, so a large window is split across parts instead of buffered whole. A window always produces exactly one manifest: even a window with zero records, which writes an empty manifest.

### The window manifest

Each window's `.manifest.json` is the authoritative commit point. It has the same schema for all three log types:

| Field | Type | Description |
| - | - | - |
| `schema_version` | string | Manifest schema version, currently `"1"` (distinct from the record `schema_version`). |
| `type` | string | Log type: `"check-run-logs"`, `"code-review"`, or `"activity-logs"`. |
| `workspace_type` | string | Your workspace type. |
| `workspace_id` | string | Your workspace identifier. |
| `window_start` | string | RFC 3339 start of the export window (UTC). |
| `window_end` | string | RFC 3339 end of the export window (UTC). |
| `exported_at` | string | RFC 3339 timestamp of when the export ran. |
| `part_count` | integer | Number of part files in this window. |
| `total_records` | integer | Total NDJSON records across all parts. |
| `parts` | array | The part files that make up this window (empty array if none). |

Each entry in `parts` contains:

| Field | Type | Description |
| - | - | - |
| `part` | integer | 1-based ordinal of the part file. |
| `file` | string | Object name of the part file, relative to the window's directory. |
| `records` | integer | Number of NDJSON records in this part. |
| `bytes` | integer | Compressed size of the part file, in bytes. |

Example manifest:

```json theme={null}
{
  "schema_version": "1",
  "type": "check-run-logs",
  "workspace_type": "github_organization",
  "workspace_id": "123456789",
  "window_start": "2026-07-16T18:00:00Z",
  "window_end": "2026-07-16T19:00:00Z",
  "exported_at": "2026-07-16T19:00:07Z",
  "part_count": 1,
  "total_records": 42,
  "parts": [
    { "part": 1, "file": "check-run-logs_20260716T180000Z_20260716T190000Z.part-0001.ndjson.gz", "records": 42, "bytes": 15734 }
  ]
}
```

### The consumer contract

<Warning>
  **Read the manifest first, then read only the parts it lists.** Any `.part-*` object under a window's directory that is **not** listed in that window's manifest is stale and must be ignored. For example, it may have been left behind by a previous, larger export of the same window. Reading part files by prefix glob without consulting the manifest can mix generations and yield inconsistent data.
</Warning>

Because the manifest is written last, a consumer that reads the manifest and then reads exactly the parts it names always sees a consistent set. Re-exporting a stable (closed) window overwrites the same part objects with byte-identical content, so a reader can never observe a torn or mixed-generation window.

### Check Run Logs

**Type:** `check-run-logs`

One NDJSON record per concluded check run in the window. Each record contains:

| Field | Type | Description |
| - | - | - |
| `schema_version` | string | Record schema version, currently `"2"`. |
| `workspace_type` | string | Your workspace type (e.g. `"github_organization"`). |
| `workspace_id` | string | Your workspace identifier. |
| `check_run_node_id` | string | GitHub node ID of the check run. |
| `check_run_type` | string | Type: `"custom"`, `"correctness"`, or `"approvability"`. |
| `custom_check_name` | string? | Name of the custom check, if applicable. Null for non-custom types. |
| `pr_number` | integer | Pull request number. |
| `repo_id` | string | Repository identifier. |
| `status` | string | Always `"completed"` (only concluded runs are exported). |
| `conclusion` | string | Check run conclusion (e.g. `"success"`, `"failure"`). |
| `created_at` | string | RFC 3339 timestamp of when the check run was created. |
| `concluded_at` | string | RFC 3339 timestamp of when the check run concluded. |
| `tool_calls` | array | Tool calls made during the check run. Empty array if none. |

Each entry in `tool_calls` contains:

| Field | Type | Description |
| - | - | - |
| `call_index` | integer | Sequential index of the call within the check run. |
| `tool_name` | string | Human-readable name of the tool used. |
| `internal_name` | string | Internal tool identifier. |
| `description` | string? | Description of what the tool call did. |
| `context` | string? | Context or reasoning for the call. |

### Code Review Analytics

**Type:** `code-review`

One NDJSON record per code review that concluded in the window. This covers both standard reviews (correctness, approvability) and per-PR custom check (CRA) reviews. Each record contains:

| Field | Type | Description |
| - | - | - |
| `schema_version` | string | Record schema version, currently `"2"`. |
| `workspace_type` | string | Your workspace type. |
| `workspace_id` | string | Your workspace identifier. |
| `repo_id` | string | Repository identifier. |
| `repo_name` | string? | Repository name. Null if unavailable. |
| `pr_number` | integer | Pull request number. |
| `pr_node_id` | string? | GitHub node ID of the PR. |
| `run_id` | string? | Review run identifier. Present for standard reviews; null for CRA per-PR records. |
| `pr_author_login` | string? | GitHub login of the PR author. |
| `pr_github_status` | string? | GitHub status of the PR (e.g. `"open"`, `"merged"`). |
| `head_sha` | string? | Head commit SHA reviewed. |
| `base_sha` | string? | Base commit SHA. |
| `head_branch` | string? | Head branch name. |
| `review_type` | string | Review type: `"correctness"`, `"approvability"`, or `"custom"`. |
| `review_started_at` | string? | RFC 3339 timestamp when the review started. |
| `review_completed_at` | string? | RFC 3339 timestamp when the review completed. |
| `cost_centimills` | integer? | Review cost, in centimills (1 centimill = 1/100,000 USD). |
| `context_bytes` | integer? | Size of the review context (diff), in bytes. |
| `incremental_context_bytes` | integer? | Size of the incremental review context, in bytes. |
| `num_objects_to_review` | integer? | Number of code objects selected for review. |
| `num_objects_reviewed` | integer? | Number of code objects actually reviewed. |
| `num_objects_billed` | integer? | Number of code objects billed. |
| `num_comments` | integer? | Number of comments produced by the review. |
| `num_fix_prs_generated` | integer? | Number of Fix It For Me PRs generated. |
| `num_fix_prs_accepted` | integer? | Number of Fix It For Me PRs accepted. |
| `num_fix_prs_not_accepted` | integer? | Number of Fix It For Me PRs generated but not accepted. |
| `any_custom_instructions` | boolean | Whether any comment in the review used custom instructions. |
| `inline_comments` | array | Inline (file/line) comments. Empty array if none. |
| `pr_level_comments` | array | PR-level (issue) comments. Empty array if none. |

Each entry in `inline_comments` contains:

| Field | Type | Description |
| - | - | - |
| `comment_id` | string | Macroscope comment identifier. |
| `github_comment_node_id` | string | GitHub node ID of the comment. |
| `github_comment_database_id` | integer? | GitHub numeric (database) ID of the comment. |
| `file_path` | string | File the comment applies to. |
| `line_start` | integer | Start line of the comment. |
| `walker` | string | Internal review identifier. It does not indicate the programming language. Older records may hold other values or an empty string. |
| `review_type` | string | Review type that produced the comment. |
| `custom_check_name` | string? | Custom check name, if applicable. |
| `has_custom_instructions` | boolean? | Whether the comment used custom instructions. |
| `comment` | string | Comment body. |
| `runtime_impact` | string? | Assessed runtime impact. |
| `pr_impact` | string? | Assessed PR impact. |
| `positive_sentiment` | boolean | Whether the comment received positive feedback. |
| `negative_sentiment` | boolean | Whether the comment received negative feedback. |
| `is_addressed` | boolean | Whether the comment was addressed. |
| `addressed_by_action` | string? | How the comment was addressed. |
| `addressed_commit_hash` | string? | Commit that addressed the comment. |
| `is_resolved` | boolean | Whether the comment thread was resolved. |
| `feedback_source` | string? | Source of the feedback signal. |
| `cost_centimills` | integer? | Comment cost, in centimills. |
| `cra_config_commit` | string? | Commit of the custom check agent config, if applicable. |
| `cra_config_file_path` | string? | Path of the custom check agent config, if applicable. |
| `created_at` | string | RFC 3339 timestamp of when the comment was created. |

Each entry in `pr_level_comments` contains:

| Field | Type | Description |
| - | - | - |
| `github_comment_node_id` | string | GitHub node ID of the comment. |
| `github_comment_database_id` | integer | GitHub numeric (database) ID of the comment. |
| `url_ref` | string? | URL reference for the comment. |
| `review_type` | string | Review type that produced the comment. |
| `custom_check_name` | string? | Custom check name, if applicable. |
| `comment` | string | Comment body. |
| `created_at` | string | RFC 3339 timestamp of when the comment was created. |

**Per-window aggregate.** In addition to the record parts and manifest, code review analytics writes one gzipped rollup object per window:

```text theme={null}
{key_prefix}{workspace_type}/{workspace_id}/code-review/{YYYY}/{MM}/{DD}/code-review_aggregate-{window_start}.json.gz
```

The aggregate is a gzip-compressed JSON object:

| Field | Type | Description |
| - | - | - |
| `schema_version` | string | Aggregate schema version, currently `"2"`. |
| `exported_at` | string | RFC 3339 timestamp of when the export ran. |
| `workspace_type` | string | Your workspace type. |
| `workspace_id` | string | Your workspace identifier. |
| `window_start` | string | RFC 3339 start of the export window (UTC). |
| `window_end` | string | RFC 3339 end of the export window (UTC). |
| `total_review_runs` | integer | Total review runs in the window. |
| `total_inline_comments` | integer | Total inline comments across all runs. |
| `total_pr_level_comments` | integer | Total PR-level comments across all runs. |
| `total_comments` | integer | Total comments (inline + PR-level). |
| `inline_positive_sentiment_count` | integer | Inline comments with positive feedback. |
| `inline_positive_sentiment_pct` | number? | Percentage of inline comments with positive feedback (0-100, one decimal). Null when there are no inline comments. |
| `inline_negative_sentiment_count` | integer | Inline comments with negative feedback. |
| `inline_negative_sentiment_pct` | number? | Percentage of inline comments with negative feedback. Null when there are no inline comments. |
| `inline_addressed_count` | integer | Inline comments that were addressed. |
| `inline_addressed_pct` | number? | Percentage of inline comments addressed. Null when there are no inline comments. |
| `inline_dismissed_count` | integer | Inline comments dismissed (resolved without being addressed). |
| `inline_dismissed_pct` | number? | Percentage of inline comments dismissed. Null when there are no inline comments. |
| `fix_it_for_me_generated` | integer | Fix It For Me PRs generated in the window. |
| `fix_it_for_me_accepted` | integer | Fix It For Me PRs accepted. |
| `fix_it_for_me_acceptance_rate` | number? | Acceptance rate of Fix It For Me PRs (0-100, one decimal). Null when none were generated. |
| `by_review_type` | array | The same metrics broken down per review type (empty array if none). |

Each entry in `by_review_type` contains `review_type` (string) plus `review_runs`, `inline_comments`, `pr_level_comments`, the four `inline_*_sentiment`/`addressed`/`dismissed` count-and-percentage pairs, `fix_it_for_me_generated`, `fix_it_for_me_accepted`, and `fix_it_for_me_acceptance_rate`: the same fields and semantics as the top-level totals, scoped to that review type.

### User Activity Logs

**Type:** `activity-logs`

One NDJSON record per workspace activity event in the window. Each record contains:

| Field | Type | Description |
| - | - | - |
| `schema_version` | string | Record schema version, currently `"2"`. |
| `event_id` | string | Unique event identifier (UUID). |
| `workspace_type` | string | Your workspace type. |
| `workspace_id` | string | Your workspace identifier. |
| `occurred_at` | string | RFC 3339 timestamp of when the event occurred. |
| `actor` | object | The actor who performed the action (see below). |
| `action_category` | string | High-level category of the action. |
| `action_type` | string | Specific action type. |
| `resource_type` | string? | Type of the resource acted on, if applicable. |
| `resource_id` | string? | Identifier of the resource acted on, if applicable. |
| `metadata` | object | Action-specific metadata. May be an empty object. |

The `actor` object contains:

| Field | Type | Description |
| - | - | - |
| `github_user_id` | integer? | GitHub numeric user ID of the actor. |
| `github_login` | string? | GitHub login of the actor. |
| `ip` | string? | Source IP address of the action. |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.