Welcome to DATTA
Your organization has too much data and too few answers: documents scattered everywhere, systems that do not talk to each other, and analyses that depend on specialists and weeks of waiting. DATTA brings intelligent search, AI-assisted analysis, dashboards and data automation together in a single platform — inside your own infrastructure — so that anyone on the team can find, understand and use information in minutes.
This is the front door of the documentation. Here you get an overview of what the platform does and the paths to go deeper into each capability. If you prefer to start with the "why", see why DATTA exists and which problems it solves.
What you can do with DATTA
Ask, instead of searching
Type a question in natural language — in the on the home page — and get an answer written by the AI together with the documents that support it, sources cited. The search understands the meaning of your question, not just the exact words. See the search guide and how to personalize the AI's answers with your own prompt.
Analyze the whole archive, not a sample
applies your compliance rules to every case in a context, classifies each one — CONFORME or ANOMALIA — and justifies the decision with cited excerpts. Details in the batch screening guide and in the audit detail of a case.
Build dashboards without waiting for the BI team
In you build interactive dashboards by dragging fields, combine different sources on the same panel and publish to your team — with a copilot that suggests visuals and narrates the results. Start with the complete DATTABI guide.
Integrate and transform data without writing plumbing
designs extraction and load pipelines visually, and the DATTAX language describes transformations, graphs and materializations in a few lines. See the Extract guide and the DATTAX language reference.
Explore data with code, when code is the right tool
The Notebook environment () offers a Jupyter integrated with the platform, with ready-made access to SQL (Trino), graphs (Neo4j, in Cypher), full-text search (OpenSearch) and PySpark — plus a
Connect your sources with governance
Every connection to a database or external source is registered in and cataloged automatically: minutes after connecting, the tables already show up in the Knowledge Catalog, with samples and quality metrics. See the connections guide.
Automate workflows
DATTA BPM models and runs business processes (screening, review, approval) with real-time monitoring. Start with the process engine guide.
First steps
A ten-minute route to feel the platform working:
- Sign in with the user your administrator provided. You land on the home page, with the search field in the center.
- Ask a real question from your day-to-day — for example, "quais processos citam prescrição intercorrente?" — and press Enter.
- Watch the screen assemble itself: the document cards appear first, and the AI's answer is written in real time just above them.
- Open the side menu and explore the sections — Descobrir, Investigação, Processar, . Every screen shows a navigation trail at the top so you never get lost.
- Ask your administrator for a workspace for your team — it groups contexts, features and panels in a single place in the menu.
How this documentation is organized
This documentation is served by the platform itself, in the Documentation Portal: open, bilingual pt-BR/EN, with instant search and previous/next navigation. The platform and operations section only appears for people with platform permission — if it is not in your menu, your profile does not reach it.
The sidebar is the map: sections group the topics, search covers the full text of every page visible to you, and the language selector switches between Portuguese and English without leaving the page.
Documentation map
User guides
For analysts, dashboard editors and pipeline operators:
- DATTABI — User Guide — DATTABI end to end (workspaces, dashboards, charts, Data Prep, Copilot, sharing, export, refresh, stream panels).
- DATTAX — Language Guide — DATTAX language (sources, transforms, full GDS catalog, AI primitives, streaming, materialization, end-to-end examples).
- DATTA Extract — User Guide — Extract Pipeline + Quick Extract (designer, components, destinations).
- Connection Management — User Guide — connection management (admin screen, inline registration, automatic cataloging, visibility).
- Appeal Co-pilot — A Real, Substantiated Appellate Brief — Appeal Co-pilot (real appellate brief substantiated with articles + case law, .docx download, Continue in Chat).
- Batch Screening — Batch Screening (live progress window, all-or-nothing, searchable report, decision + appeal admissibility).
- Document Upload and Ingestion — upload and ingestion (mandatory target context, real-time feed of recognized entities, post-refresh resumption).
- Case Audit Detail — audit detail (Detected Anomaly, Case Summary, rules without context).
Administration guides
For platform administrators:
- JDBC Drivers — Official Catalog and Governance — JDBC drivers (official catalog of 43 drivers, initial load, governance, manual download).
- Managing Contexts — the Sistema → Contextos screen (create/edit/update/clear/delete a context, per-feature Visibility Matrix).
- Multi-source context — multi-source context (dataSources) + procedural code base (
codigoProcessualDatabase). - Complete Execution History — complete execution history (durable storage, filter by period, orphan task reconciliation).
Operational runbooks
- Runbook — dattabi-pool (node pool) — the
dattabi-poolnode pool (cloud and on-premises provisioning, quotas, automatic scaling, network, troubleshooting). - Runbook — Streaming (F9) — Kafka (Strimzi), Debezium Connect, Pulsar, Neo4j CDC, periodic OpenSearch reads and the dedicated streaming pool.
- Runbook — Connection Management — operating the central connection registry (deployment, encryption-key rotation, tuning of automatic cataloging, backup/restore, auditing, orphan connections).
- Runbook — Incident Response — playbooks for common incidents (Kafka down, Neo4j out of memory, cache unavailable, DATTAX engine overloaded, blocked driver scan, slow OpenSearch, 5xx errors at the platform's front door).
- Runbook — Spark Operations — Spark (submission, cancellation, logs, native UI, scaling, troubleshooting).
Spark administration
- Spark Admin Console (Phase 4.6) — overview of Spark administration (four aggregated sources: cluster resources, driver API, History Server, metrics).
- Spark Admin — Architecture — architecture, two-tier cache, resource watcher, reverse proxy.
- Spark History Server — deploying the Spark History Server and configuring event logs.
Infrastructure console
- Infrastructure Console — DATTA's infrastructure console (applications, configurations, credentials, volumes, jobs and scheduled jobs, events, nodes, automatic scaling, operators).
- Infrastructure Console — Architecture — collection, metrics and cache; event watcher; destructive-mutation flow.
- Infrastructure Console — RBAC and auditing — the PLATFORMVIEW / PLATFORMADMIN / PLATFORM_SENSITIVE tiers, rate limits, audit event format.
- Runbook — Platform Operations — the mapping between command-line operations and the console (compliance reference for operating everything through the interface).
Cluster and node pools
- Cluster Management (Phase 4.9) — overview of cluster management (node pools, host onboarding, control plane, certificates, etcd, maintenance, declarative provisioning).
- Node Pools — the pool concept, seed defaults, batch operations, resource quotas.
- Host Registration (Join) — host onboarding per distribution, SSH-assisted, pre-checks and troubleshooting.
- Certificate Rotation — automatic monitor + rotation job + rollback procedure.
- etcd Operations — snapshot, restore, health and destination storage.
- CAPI Integration (optional) — runtime detection, machine scaling, on-premises installation.
- SSH hardening — Phase 4.10 — embedded SSH client, AES-GCM vault for credentials, TOFU known-hosts, rotate/revoke/list.
- Upgrade Automation — Phase 4.10 — seven-phase state machine, preflight and postflight checks, rolling upgrade, honest rollback, approval gates.
- Cluster Operations Runbook — the console equivalent of more than 40 command-line operations.
- Runbook — Cluster upgrade — upgrade walkthrough (prerequisites, gates, rollback, verification).
- Node Maintenance — node maintenance workflows (patching, agent upgrade, rolling drain).
Day-zero installation
- Day-0 Installer (DATTA) — day-zero overview.
- ClusterBlueprint — fields — installation blueprint fields.
- Distro matrix — Phase 4.11 — distribution × operating system matrix.
- CNI options — Phase 4.11 — cluster networking (Cilium by default, Calico/Flannel as alternatives).
- Air-gapped installation (no internet) — HTTP proxy and offline registry for environments without internet access.
- Runbook — DATTA day-zero install — installation from zero, start to finish.
- Runbook — Enabling the ISO builder (Packer) — enabling installation image generation.
Data-layer subconsoles
- Data Layer Console — five subconsoles (Neo4j, OpenSearch, platform cache, Kafka and AI) consolidated in one place.
- Data Layer Console — Architecture — internal organization, cache, three-tier RBAC, audit events and delegation to the infrastructure console.
- Runbook — Neo4j Operations — database lifecycle, cluster, indexes, users, read-only Cypher, backup/restore.
- Runbook — OpenSearch Operations — cluster health, indices, lifecycle policies, query tools, snapshots, dashboards.
- Runbook — Platform cache operations — the platform cache: six-node cluster, key explorer (with masked secrets), failover, command statistics.
- Runbook — Kafka Operations — topics, consumer groups and offset resets, message preview, Debezium connectors, schema registry.
- Runbook — LLM / vLLM Operations — AI models, providers (local or external), usage and costs, catalog.
Architecture
- Architecture — Overview — overview, components and main flows.
- Architecture — Data Access — the three data-access layers (drivers → connection registry → interactive and bulk execution), JDBC vs. native connector matrix and the automatic cataloging flow.
- Architecture — DATTABI Data Mesh — materialization planner, destinations (Iceberg/Neo4j/OpenSearch), refresh policies and lineage.
- Architecture — Copilot and LLM — chat and model management, Copilot at edit time vs. view time, narrative, RAG and the DATTAX AI primitives.
- Architecture — Streaming — feasibility per source, enrichment, windowing, state, WebSocket/SSE.
- Architecture — Security — RBAC matrix, credential vault, audit event catalog, rate limits, circuit breakers and guards against improper access.
- End-to-end Neo4j database name consistency — canonical database name end to end, the target database is never guessed (ingestion) and dynamic resolution of the context database (Project Guidelines §9).
- Architecture — Semantic Text Search over Chat History — semantic text search (KNN+BM25 via RRF) over chat history, scoped per user.
- NER/Gemini resilience — 429, throttling and key rotation — resilience of entity recognition (retries with progressive backoff, global throttling, key rotation).
- Screening Retrieval — generic per-context identifier — screening retrieval by generic identifier (KNN+BM25 via RRF), resolution through the context configuration, case↔excerpt linkage and the durable ingestion fix.
GA validation
- GA Checklist — Project Guidelines Audit — Project Guidelines audit, item by item.
- Security Review — Pre-GA — security review (injection, SSRF, improper access, credentials, rate limits, auditing).
- Architecture Review — Pre-GA — architecture review (single responsibility, coupling, scalability, observability, disaster recovery).
- Known Issues — Pre-GA — post-GA items with severity.
- Review Questions — Pre-GA Validation — review questions to submit before GA.
DATTAX reference
- DATTAX — Graph Algorithm Catalog (Neo4j GDS) — the complete catalog of Neo4j GDS graph algorithms (more than 30).
- DATTAX Streaming — Syntax — streaming syntax.
Infrastructure (summaries)
- DATTA Node Pools — node pool summary.
- JDBC Drivers — Official Catalog and Governance — JDBC driver management.
- Architecture — Streaming — streaming and CDC.
Datta Intelligent Query Layer (formerly SIQL)
- Datta Intelligent Query Layer — Discovery and Foundations — Phase 0 (mapping the current stack).
- datta-intelligent-query-layer — Architecture (Phase 1 MVP) — technical design, contracts and knowledge graph.
- DATTA Intelligent Query Layer — Phase 2 Architecture — Phase 2 (telemetry + observed + first inferred).
- DATTA Intelligent Query Layer — Phase 3 Architecture (Learning Layer) — Phase 3 (offline/nearline learning + GDS + feature store).
- SIQL — Phase 4 — Physical Maintenance — Phase 4 (physical Iceberg maintenance).
- SIQL — Phase 5: Semantic + Ontology + vLLM Assistance — Phase 5: ontology + AI assistance + NL-to-SQL.
- SIQL — LLM Usage Boundaries — critical: boundaries of AI usage inside SIQL.
- SIQL — NL-to-SQL — NL-to-SQL design and safety rails.
- datta-intelligent-query-layer — Operational Runbook — deployment, rollback, metrics, alerts and troubleshooting.
- SIQL Phase 2 — Benchmark Plan — GA criteria per phase (including Phase 5 targets).
- datta-intelligent-query-layer — Apache Trino Integration — how to install the plugin on an existing Trino.
- SIQL Trino Admin Console (Phase 4.5) — Trino administration console.
- Configuring per-table maintenance policies — physical maintenance policies.
- SIQL — Model Registry — registry of learned models.
Other
- Configuration: Trino 480 (Replaces Impala) — Trino configuration.
- Guide: DATTA Multi-Source Notebook and DATTA Notebook — documentation and examples — the Notebook environment.
- Data Catalog + Ontology Architecture -- DATTA Platform — catalog and ontology.
- ETL/Pipeline Architecture -- DATTA Platform — ETL pipeline.
- Vendoring the frontend runtime libraries (§6/§15/§17) — frontend libraries served by the platform itself.
- Guide: vLLM as an Optional Provider (Air-gapped / Cost) — migration of the local inference engine.
For developers and integrations
Practically everything the interface does is also available through the API — see the DATTA API overview and the endpoint reference.