# Confluence Connector - FAQ

## General

### What type of connector is this?
**Answer:** The Confluence Connector is a pull-based synchronization service that periodically scans Confluence spaces for labeled pages and syncs their content and attachments to the Unique knowledge base.

**Key characteristics:**
- Runs on a configurable cron schedule (default: every 15 minutes)
- Pulls content from Confluence via REST API
- Requires explicit labeling of pages to trigger synchronization
- Operates as a background service without user interaction
- Supports both Confluence Cloud and Confluence Data Center

### How does this differ from the Confluence Connector v1?
**Answer:**
| Aspect | v1 | v2 |
| --- | --- | --- |
| Multi-tenancy | Not supported | Multiple Confluence instances in a single connector pod |
| Attachment ingestion | Not supported | Supported with configurable MIME types and size limits; embedded images are inlined into their page artifact |
| Change detection | File-diff mechanism (pages only) | File-diff mechanism (pages and attachments) |
| Safety guards | None | Full-deletion prevention, concurrent sync prevention |
| Key format | `spaceId_spaceKey/pageId` | `tenantName/spaceId_spaceKey/pageId` (v1 format available via `useV1KeyFormat`) |

## Labels and Page Discovery

### How does the connector decide which pages to sync?
**Answer:** The connector uses two configurable Confluence labels: one for single-page sync (e.g. `ai-ingest`) and one for syncing a page and all its descendants (e.g. `ai-ingest-all`). Both labels must be explicitly set in the tenant configuration.

### What happens when a page has both labels?
**Answer:** The page is deduplicated. The connector merges all labeled pages and their descendants into a single unique set (by page ID), so no page is ingested twice.

### What happens when a page has `ai-ingest` and its ancestor has `ai-ingest-all`?
**Answer:** The page is discovered through both paths but deduplicated by ID, so it is ingested exactly once.

### Which Confluence content types are synced?
**Answer:** The connector ingests `page` and `blogpost` content (including Live Docs, which are a page subtype). Attachments are ingested conditionally when `attachments.mode=enabled`. Content types `database`, `whiteboard`, and `embed` are explicitly skipped because their APIs expose no renderable body.

### What format is the page content exported in?
**Answer:** Pages are fetched using the `body.storage` expansion, which returns the Confluence storage representation (HTML). The content is uploaded to Unique with MIME type `text/html`.

### Are Confluence labels preserved during ingestion?
**Answer:** Yes. All labels on a page are included as metadata during ingestion, except for the two connector labels (`ai-ingest` and `ai-ingest-all` by default), which are filtered out.

### Which spaces are scanned?
**Answer:** Only `global` spaces are scanned (Cloud also includes `collaboration` spaces). Personal spaces are excluded on both platforms.

## Authentication

### What authentication methods are supported?
**Answer:** The connector supports OAuth 2.0 (2LO) for Confluence Cloud and Data Center (10.1+), which is the recommended authentication method. Personal Access Token (PAT) is supported only on Data Center versions below 10.1 where OAuth 2.0 (2LO) is not available, and is not recommended.

### How are secrets managed in configuration?
**Answer:** Secret values in tenant YAML configuration files use the `os.environ/VARIABLE_NAME` syntax to reference environment variables, resolved at startup.

## Configuration

### How are tenants configured?
**Answer:** Each tenant is configured via a YAML file following the naming convention `<tenant-name>-tenant-config.yaml`. Tenant names must match the pattern `^[a-z0-9]+(-[a-z0-9]+)*$` and must be unique across all config files.

### What are the available tenant statuses?
**Answer:**
| Status | Behavior |
| --- | --- |
| `active` | Tenant is loaded and sync is scheduled (default if not specified) |
| `inactive` | Tenant config is validated but the tenant is not loaded |
| `deleted` | Ingested content is deleted from the Unique knowledge base and sync is stopped |

### What are the key configuration sections?
**Answer:** Each tenant YAML file contains four top-level sections:

| Section | Purpose |
| --- | --- |
| `confluence` | Instance type, base URL, authentication, API rate limit, label names |
| `unique` | Unique API endpoints, authentication mode, rate limit |
| `processing` | Concurrency, cron schedule, optional scan limit |
| `ingestion` | Ingestion mode, scope ID, attachment settings, v1 key format toggle |

### What are the default values for key settings?
**Answer:**
| Setting | Default Value |
| --- | --- |
| Processing concurrency | 1 |
| Scan interval cron | `*/15 * * * *` (every 15 minutes) |
| Unique API rate limit | 100 requests/minute |
| Attachment ingestion | Enabled |
| Maximum attachment size | 200 MB |
| Store internally | Enabled |
| Use v1 key format | Disabled |

### What attachment MIME types are allowed by default?
**Answer:** The defaults cover PDF (`application/pdf`), the major Office formats (DOCX, XLSX, PPTX), plain text, CSV, HTML, PNG, and JPEG. These can be overridden via `ingestion.attachments.allowedMimeTypes`.

### Are embedded images ingested?
**Answer:** Yes. Each referenced PNG or JPEG attachment is inlined into the page HTML as a base64 data URI.

### What happens if an image cannot be inlined into its page?
**Answer:** When inlining is enabled, image attachments are not ingested as standalone artifacts.

### What happens to image attachments that were ingested as separate artifacts before inlining was introduced?
**Answer:** They are not automatically removed. The file-diff mechanism only deletes content whose discovery key disappears from Confluence.

### How do I verify that page image inlining is working after a deployment?
**Answer:** Pick a Confluence page known to contain an inline image, run a sync, and inspect the ingested page artifact in the destination scope.

### How do I find my Atlassian Cloud ID?
**Answer:** The Cloud ID is required only for Confluence Cloud instances: `https://your-domain.atlassian.net/_edge/tenant_info`

## Sync Behavior

### What happens during a sync cycle?
**Answer:** Each sync cycle follows these steps:
1. Grant the service account access to the pre-existing root scope in Unique.
2. Discover all pages matching the configured labels via CQL search
3. Fetch descendant pages for any pages with the all-descendants label
4. Extract allowed attachments from discovered pages
5. Compute a file diff per space against Unique's stored state
6. Create child scopes in Unique for each space (using the space key as scope name)
7. Fetch and ingest new or updated pages (HTML storage representation)
8. Download and ingest new or updated attachments (streamed)
9. Delete items from Unique that are no longer discovered
10. Detect space scopes whose Confluence space is no longer discovered, and remove their files and scopes

### How does change detection work?
**Answer:** The connector uses a server-side file diff mechanism that compares discovered items per space against the state stored in Unique.

### What happens when a label is removed from a page?
**Answer:** If the `ai-ingest` label is removed from a page, the page is no longer discovered during the scan.

### What happens when a page is deleted from Confluence?
**Answer:** If the page's space is still discovered during the next sync cycle, the file diff detects the missing page and deletes the corresponding content from Unique.

### What happens to attachments when their parent page is unlabeled?
**Answer:** Attachments are discovered as children of labeled pages. If a page is no longer discovered, its attachments are also missing from the discovery results.

### How are scopes organized in Unique?
**Answer:** Scopes follow a two-level hierarchy: a root scope configured per tenant, and child scopes automatically created for each Confluence space key.

## Safety and Deletion

### What safety guards does the connector have?
**Answer:** The connector includes the following safeguards to prevent accidental data loss:
- **Zero-submission guard**: Sync cycle aborted if zero items are discovered for a space.
- **Full-deletion guard**: Sync cycle aborted if the file diff indicates deletion of every file for a space.
- **Root scope ownership validation**: Prevents sync for a tenant if the root scope is owned by a different Confluence instance.

### What happens if I reassign the root scope to a different Confluence instance?
**Answer:** This is not supported. The sync will abort with a fatal error.

### Are concurrent syncs for the same tenant possible?
**Answer:** No. If a sync cycle is already running, the new cycle is skipped.

## Troubleshooting

### Why aren't my pages syncing?
**Checklist:**
1. Does the page have the `ai-ingest` or `ai-ingest-all` label?
2. Does the service account have access to the page's space?
3. Is the page in a global space?
4. Is the page a standard page type?
5. Does the page have a non-empty body?
6. Is the tenant status set to `active` in the YAML config?
7. Check connector logs for errors.

### Why aren't attachments being ingested?
**Checklist:**
1. Is attachment ingestion enabled?
2. Does the attachment's `mediaType` appear in the `allowedMimeTypes` list?
3. Is the file at most the configured `maxFileSizeMb`?
4. Is the file size greater than 0 bytes?
5. If the attachment is an image and chunks are missing, check that `attachments.imageOcr` is `enabled`.
6. Check connector logs for attachment-specific errors.

### Why do I see "Aborting to prevent accidental full deletion" errors?
**Answer:** The full-deletion safety guard was triggered.

**Resolution:** Check Confluence API connectivity and authentication.

### Why is sync taking too long?
**Possible causes:**
- Large number of labeled pages and descendants
- Large attachments being downloaded and uploaded
- Low API rate limit configuration
- Low processing concurrency

**Solutions:**
1. Increase `processing.concurrency`
2. Increase `confluence.apiRateLimitPerMinute`
3. Increase `unique.apiRateLimitPerMinute`
4. Review labeled pages and reduce scope if necessary
5. Adjust `processing.scanIntervalCron` to allow more time between cycles.

### How does the connector handle errors during ingestion?
**Answer:** Individual item failures are logged and skipped without aborting the entire sync cycle.

## Multi-Tenancy

### Can one connector serve multiple Confluence instances?
**Answer:** Yes. Each Confluence instance is configured as a separate tenant with its own YAML configuration file.

### How do I add a new tenant?
**Answer:** Create a new YAML configuration file following the naming convention `<tenant-name>-tenant-config.yaml`.

### Can two tenants use the same scope ID?
**Answer:** No. Each root scope is tagged with the tenant that owns it.

## Performance

### What are the API rate limits?
**Answer:** Both Confluence and Unique API rate limits are independently configurable per tenant:

| API | Configuration Key | Default |
| --- | --- | --- |
| Confluence | `confluence.apiRateLimitPerMinute` | No default (must be set) |
| Unique | `unique.apiRateLimitPerMinute` | 100 requests/minute |

### What is the initial sync behavior?
**Answer:** An initial sync is triggered immediately on startup for each active tenant.

## Permissions

### What Confluence permissions does the connector need?
**Answer:** The connector requires read access to the Confluence instance.

### What happens if the connector lacks permission to a space?
**Answer:** Pages in inaccessible spaces are silently excluded from CQL search results.

### How does Unique platform authentication work?
**Answer:** The connector supports two modes: `cluster_local` for in-cluster deployments and `external` for out-of-cluster deployments.

## Resource Requirements

### What are the resource requirements?
**Answer:** The default Helm chart values specify the following Kubernetes resource settings:

| Resource | Value |
| --- | --- |
| CPU request | 1 core |
| Memory request | 512 Mi |

These defaults are suitable for a single-tenant deployment with moderate page counts.
