Original data · Quarterly snapshot

AIBookCraft Public Story Pulse

A privacy-preserving view of what sits on the public AIBookCraft bookshelf: genres, languages, story size, and recurring writing attributes across 2,593 published stories.

Snapshot: Read the methodology

The snapshot at a glance

2,593

public, published stories

13,881

median words per story

10

median episodes per story

9

languages represented

Descriptive findings

Three things visible in this catalog

  • Fantasy is the largest named primary-genre bucket.

    It accounts for 413 stories (15.9%). The separate “Other / unclassified” bucket includes unsupported and missing labels.

  • The middle half spans 9,04025,712 words.

    Episode counts run from 6 to 20 across the same 25th–75th percentile band. Zero values remain in the calculation.

  • English metadata appears on 82.8% of stories.

    The catalog contains 9 language labels in total. This reflects metadata, not validated linguistic analysis of the prose.

Catalog composition

Primary genre distribution

Every story contributes to one controlled primary-genre bucket. Unsupported or missing values are grouped rather than exposed as raw labels.

Primary genre distribution in the public story catalog
CategoryStoriesShare of catalog
Other / unclassified470
18.1%
Fantasy413
15.9%
Adventure243
9.4%
Non-fiction220
8.5%
Fiction216
8.3%
Mystery216
8.3%
Romance166
6.4%
Historical143
5.5%
Children's117
4.5%
Fairy tale97
3.7%
Science fiction85
3.3%
Comedy66
2.5%
Poetry64
2.5%
Essay38
1.5%
Horror15
0.6%
Drama12
0.5%
Thriller7
0.3%
Biography3
0.1%
Autobiography2
0.1%

Publishing languages

Language distribution

These are the language values attached to catalog records. They have not been independently detected from the story text.

Language distribution in the public story catalog
CategoryStoriesShare of catalog
English2,146
82.8%
French334
12.9%
Korean45
1.7%
Arabic30
1.2%
German19
0.7%
Chinese6
0.2%
Japanese6
0.2%
Russian6
0.2%
Spanish1
0.0%

Story shape

Word and episode count percentiles

Percentiles describe the spread without letting a few very long books define the “typical” story. All 2,593 finite, non-negative catalog values—including zero—are included.

Word-count and episode-count percentiles for public stories
MeasureMinimum25thMedian75th90thMaximum
Total word count(words)09,04013,88125,71233,197225,828
Episode count(episodes)06102020464

Recurring attributes

Common public story tags

A tag appears here only when it is attached to at least five stories. Percentages use the full public catalog as the denominator, and a story may carry multiple tags.

Common privacy-thresholded public story tags
CategoryTagged storiesShare of catalog
Pace: Normal2,166
83.5%
POV: Third Person2,053
79.2%
Audience: General1,977
76.2%
Tone: Warm1,087
41.9%
Tone: Adventure562
21.7%
Tone: Mystery543
20.9%
POV: First Person530
20.4%
Audience: Children410
15.8%
Ages 7–9377
14.5%
Pace: Slow231
8.9%
Tone: Funny201
7.8%
Audience: Adult196
7.6%
Tone: Poetic188
7.3%
Pace: Fast186
7.2%
Romance10
0.4%

Use the snapshot as a starting point, not a prescription

A catalog can show the shelves that already exist. It cannot tell you which story only you can write. Explore examples, then use a focused tool to turn your own premise into a plan.

Reproducible research

Methodology and limitations

Unit of analysis
One public, published, moderation-visible story in the AIBookCraft catalog at extraction time.
Categories
Each story contributes once to its primary genre and language. Genre values are mapped to the site's controlled public genre taxonomy; unmatched or missing values are grouped as Other / unclassified.
Tags
Tags are normalized case-insensitively and counted at most once per story. Internal builder/context markers and contact-like or unrecognized structured strings are excluded. A tag is shown only when used by at least 5 stories, with at most 24 reported.
Percentiles
Word and episode summaries include finite, non-negative values, including zero. Percentiles use linear interpolation on the sorted catalog values and are rounded to whole numbers.
Privacy
The published snapshot contains no story IDs, titles, descriptions, cover URLs, prompts, author identifiers, or raw story text.

Limitations

  • This describes the composition of the public catalog, not reader demand, search volume, story quality, or commercial performance.
  • Genre, language, tag, episode, and word-count metadata are creator-supplied or product-generated and may be missing or inconsistent.
  • The catalog can change while paginated extraction runs, so the result should be treated as a dated snapshot rather than a transactional count.
  • A single snapshot cannot establish a trend or a causal relationship. Comparable repeated snapshots are needed for change-over-time claims.