Oak Index Studio

/oak:index — query → lucene definition

TargetGuide

/analysis

Node type

cq:Page

Path restriction

/content/wknd

order: jcr:content/cq:lastModified ↓

/analysis/health

Overall score

100/100

Excellent
  • includedPaths (index root) ✓

    includedPaths (/content/wknd) scopes the index to specific content — good.

  • queryPaths (index root) ✓

    queryPaths matches includedPaths — good.

  • ordered (all properties) ✓

    ordered=true properties detected — this alone is sufficient for ORDER BY (doc-values sorting); propertyIndex is only needed if the property is also filtered in a WHERE clause.

  • property types (all properties) ✓

    No obviously mistyped properties detected.

  • propertyIndex (all properties) ✓

    Every property definition has at least one capability flag set — good.

  • analyzed (none) ✓

    No analyzed properties.

  • nodeScopeIndex (all properties) ✓

    nodeScopeIndex usage is consistent — always paired with analyzed.

  • evaluatePathRestrictions (index root) ✓

    evaluatePathRestrictions=true — good.

  • null checks (all properties) ✓

    No nullCheckEnabled properties — good.

  • regex (all properties) ✓

    No regex-based property definitions — good.

  • tags (index root) ✓

    No index tag set — fine unless Oak is currently selecting the wrong index for this query.

  • selectionPolicy (index root) ✓

    No selectionPolicy set — Oak's default cost-based selection applies.

  • compatVersion (index root) ✓

    compatVersion=2 — good.

  • async (index root) ✓

    async=["async","nrt"] — async with near-real-time visibility, good.

/analysis/performance-estimate

Estimated complexity

Low(29/100 heuristic)

Relative cost heuristic (arbitrary scale — not Oak's cost=X)

~14–41

Confidence

65/100

Reduced from a 70-point ceiling (this tool never claims high confidence) for: 1 range predicate(s) of unknown width.

  • Index selectivity [medium]

    Best-case selectivity across 2 properties: moderate. Oak evaluates ANDed conditions by intersecting, so the single most selective predicate dominates.

    Assumption: Selectivity is derived only from operator type (=, range, LIKE, IS NULL, ...) combined with the guessed cardinality above — never from actual value distribution.

  • Property cardinality assumptions [low]

    jcr:content/cq:template: low (~10-50 values assumed); jcr:content/cq:lastModified: high (near-continuous)

    Assumption: Cardinality is guessed from property name and declared type only (Boolean → 2 values, Date → near-continuous, enum-like names → ~10-50 values, everything else defaults to an unverified 'medium' guess).

  • Range predicate impact [medium]

    1 range predicate(s) on jcr:content/cq:lastModified. Cost scales with how much of the value space each interval covers.

    Assumption: This parser can't tell a tightly-bounded range (e.g. a single day) from a wide-open one (e.g. 'after 2020') from the query text — assumed 'moderate, unknown width' for every range predicate found.

  • Join cost [low]

    No JOIN — zero join cost.

    Assumption: N/A — single-selector query.

  • Sorting cost [low]

    ORDER BY on 1 property. A single ordered property can be pre-sorted by lucene cheaply; multiple properties typically fall back to an in-memory sort of the full result set.

    Assumption: Assumes the generated index defines ordered=true on the sorted properties (as this app's own generator does) — if it doesn't, sorting cost is worse than estimated here.

  • Path restriction effectiveness [medium]

    Scoped path(s) (/content/wknd, depth 2). Deeper paths are assumed to narrow the candidate set more.

    Assumption: Heuristic only: path depth is used as a rough proxy for 'how much of the repository this excludes' — this tool has no actual knowledge of how content is distributed under that path.

Assumptions behind this estimate

  • This estimator has no access to the real repository, its content, or Oak's actual index statistics — every number below is a relative heuristic for comparing queries, not a prediction of real query time.
  • Property cardinality (how many distinct values a property has) is guessed from its name and declared type, never measured against real data.
  • Range predicate width (how much of the value space a >, <, or BETWEEN-style condition actually covers) can't be determined from the query text alone.
  • Join cost assumes Oak's documented per-selector-independent-then-merge execution model; the real cost depends on each selector's actual result-set size, which is unknown here.
  • jcr:content/cq:template's name suggests an enum-like field (status/type/template/tag/...) — assumed roughly 10-50 distinct values, not measured.
  • jcr:content/cq:lastModified is a Date — assumed near-continuous distinct values, so exact-value equality is expected to be highly selective (range comparisons are assessed separately).

All figures on this page are heuristic estimates derived from the query's shape alone, on an arbitrary 0-100 scale — they are NOT Oak's real cost() values, and are not expressed in Oak's costPerEntry/costPerExecution units. Oak's actual cost is computed from real entryCount statistics in the live repository (cost ≈ entryCount × costPerEntry + costPerExecution), which this tool has no access to and therefore cannot emulate. Always verify with an actual Explain Query (cost=X) against the target repository before drawing conclusions.

/analysis/properties

propertyopstypeflags
jcr:content/cq:template=String
jcr:content/cq:lastModifiedrange,orderDateordered

/analysis/reasoning

  • index root type=lucene / compatVersion=2

    Lucene compat 2 is the only index type supporting the combination of property, ordered, analyzed and facet definitions in one index; required on AEMaaCS.

  • index root async=["async","nrt"]

    Async + NRT: near-real-time updates between async cycles, standard for AEMaaCS.

  • index root includedPaths=["/content/wknd"]

    Only content under the query's path restriction is indexed — smaller index, faster reindex.

  • index root evaluatePathRestrictions=true

    Stores :ancestors so ISDESCENDANTNODE / path= is evaluated inside lucene instead of post-filtering every hit.

  • index root queryPaths

    Tells the query engine this index only answers queries under these paths — prevents wrong index selection for unrelated queries.

  • jcr:content/cq:template propertyIndex=true

    Query filters on this property (=) — without propertyIndex the value is not stored for lookup and the condition would be post-filtered.

  • jcr:content/cq:lastModified propertyIndex=true

    Query filters on this property (range) — without propertyIndex the value is not stored for lookup and the condition would be post-filtered.

  • jcr:content/cq:lastModified ordered=true

    Range comparison (>, <, BETWEEN, daterange) needs ordered storage to seek instead of scanning all values.

  • jcr:content/cq:lastModified type=Date

    Value/format in the query implies Date; typed storage makes range comparison and ordering correct (string-ordered dates/numbers sort wrongly).

  • index root indexRules/cq:Page

    Rule scoped to cq:Page — only nodes of this type (and subtypes) are indexed, keeping the index minimal.

/analysis/selectors

Selector

p

Type

cq:Page

Path restriction

/content/wknd

Properties

p

  • jcr:content/cq:template
  • jcr:content/cq:lastModified

Order by

p.jcr:content/cq:lastModified ↓

/oak:index/cqPageLucene-custom-1

/warnings

No warnings.

/suggestions

  • cq:Page already has a substantial OOTB Lucene index (/oak:index/cqPageLucene). Before deploying this generated definition as a wholly separate index, check whether your project already has a cqPageLucene-custom-N copy and add these property definitions there instead — two independent indexes covering the same node type roughly doubles write-time indexing overhead for every cq:Page change. Never edit /oak:index/cqPageLucene itself (product updates can overwrite it) — a -custom-N copy, which is exactly the naming this tool already generates on AEMaaCS, is the correct way to extend it.
  • AEMaaCS naming: 'cqPageLucene-custom-1' follows the required <name>-custom-<version> convention; deploy under /oak:index in ui.apps — Cloud Manager triggers reindexing automatically. Never set reindex=true in the package on AEMaaCS.
  • Deliberately omitted: tags, selectionPolicy, costPerEntry/costPerExecution. Add an index tag + option(index tag ...) only if Oak picks the wrong index; cost overrides are a last resort and mask real problems.

/performance

Query today (no new index)
45/100
With generated index
99/100

Heuristic estimate from static analysis. Real cost depends on repository size and index statistics — verify with Explain Query after deployment.

AEM OAK LUCENE REFERENCE & BEST PRACTICES

Understanding AEM Oak Lucene Indexing & Query Tuning

Apache Jackrabbit Oak powers the content repository across Adobe Experience Manager (AEMaaCS, AEM 6.5, and AMS). Because Oak is a hierarchical content repository, unindexed JCR-SQL2, XPath, and Query Builder queries cause full content traversals. Oak Index Studio generates production-grade oak:QueryIndexDefinition nodes designed to prevent traversals and maximize query throughput.

01 / DETERMINISTIC RULE ENGINE

Automatically determines propertyIndex, ordered, type, and compatVersion flags directly from your query constraints without guesswork or AI hallucinations.

02 / 14 INDEX HEALTH CHECKS

Validates path scoping (includedPaths / queryPaths), async modes (async / nrt), null-check definitions, and warns against duplicate OOTB indexing.

03 / EXPLAIN & XML DIFFING

Inspect Oak Explain plan cost outputs to diagnose index selection decisions, or compare your existing .content.xml against query requirements.

Frequently Asked Questions (FAQ)

Common questions on AEM index creation, query performance, and deployment.

Why are my AEM queries failing with QueryTraversalException in production?

In Apache Jackrabbit Oak, when a query executes without a suitable Oak Lucene index, Oak must traverse nodes hierarchically under the query path. To prevent repository degradation, Oak halts the query when traversal exceeds the safety threshold (defaulting to 10,000 or 100,000 nodes) with a QueryTraversalException. Generating a Lucene index containing property definitions for your filtered properties eliminates traversal.

What happened to OakUtils, and is Oak Index Studio a replacement?

OakUtils (formerly hosted at oakutils.appspot.com) was the most popular community tool for generating Oak Lucene index XMLs from AEM queries before becoming permanently unreachable. Oak Index Studio was designed as a modern, open-source replacement with zero server dependencies, instant real-time analysis, index health evaluations, heuristic cost scoring, and Explain plan inspection.

How does Oak Index Studio compare to generating indexes with ChatGPT or generic AI?

General LLMs frequently hallucinate or misapply Oak constraints. For instance, LLMs often assign propertyIndex=true to negated LIKE clauses (which Oak cannot index), hallucinate unnecessary evaluatePathRestrictions=true, or enforce propertyIndex alongside ordered=true for sorting-only properties. Oak Index Studio uses deterministic, unit-tested rules based directly on official Apache Jackrabbit Oak specifications.

How do I deploy generated Oak index definitions on AEM as a Cloud Service (AEMaaCS)?

On AEMaaCS, custom indexes extending out-of-the-box indexes must follow the naming convention <indexName>-custom-<version> (such as cqPageLucene-custom-1). Deploy the definition under /oak:index inside your project's ui.apps module. Never set reindex=true in Cloud Service packages — Cloud Manager automatically detects the version increment and handles background reindexing during deployment.

When should I use propertyIndex=true vs ordered=true in an Oak Lucene index?

Use propertyIndex=true when a property appears in query filter criteria (equality '=', IN(), or bounded LIKE) so Oak indexes the value for rapid lookup. Use ordered=true when a property is used in an ORDER BY clause or in inequality/range comparisons (>, <, BETWEEN, daterange) to enable Lucene DocValues seek instead of in-memory sorting.

Is my query text or repository data sent to any remote server?

No. Oak Index Studio operates entirely client-side inside your browser tab. No server APIs, telemetry, or external trackers are contacted. Your repository structure, property names, and query details never leave your local environment.