What Makes a Poem Great? A 10-Criterion Test

What makes a poem great — Poetic Excellence Index by Danil Rudoy

Can poetic greatness be measured? The Poetic Excellence Index tests artistic execution through textual evidence, blinded expert judgment and reader response.

People disagree about poems for good reasons: taste, training, temperament, culture, memory, and mood all change what a reader notices and values. That variation still leaves testable questions. A poem has observable features, and groups of readers can be studied for reproducible responses to them.

Quick answer. A great poem solves the artistic problem it creates for itself with exceptional precision, coherence, necessity and force. Different poems activate different resources: an epigram may depend on compression and wit, an elegy on voice and affect, a narrative poem on development, and an imagist lyric on perceptual exactness. The PEI therefore combines ten common dimensions with a pre-scored task profile that identifies which dimensions are central, supporting or genuinely inapplicable.

Subjectivity creates variance; it does not erase signal.

This article proposes a falsifiable evaluation framework for poetic excellence. A numerical result functions as a structured claim open to testing and revision. The governing hypothesis is that poetic greatness contains a reproducible evaluative signal beyond private taste because many relevant textual properties are observable and can be independently rated.

The framework has two separate outputs. The Poetic Excellence Index (PEI) rates the artistic performance of a poem on ten criteria, for a maximum of 100 points. The Canonical Anchor Index (CAI) estimates a different property: how strongly a poem may function as a durable cultural anchor. PEI excludes fame. Historical influence belongs to evidence accumulated through time.

Three kinds of evidence

A serious judgment of poetry should keep three levels of evidence distinct.

1. Observable properties of the text

Meter, rhyme, syntax, lineation, repetition, phonetic patterning, parallelism, deviation, stanza architecture, lexical recurrence, image systems and structural turns can be identified in the text. Quality depends on execution, interaction and necessity. A regular meter may strengthen a poem, constrain it productively, or force its diction. Free verse can display equally rigorous control through cadence, syntax, line break and patterned recurrence. For related formal vocabulary, see rhyming poetry, rhyme schemes and meter.

2. Intersubjective evidence

Independent judges can disagree while still showing meaningful agreement. Teresa Amabile’s Consensual Assessment Technique operationalized creativity through independent judgments by people familiar with a domain [1], and later work has examined how expertise changes the reliability of those judgments [2]. In poetry, this suggests a testable question: when qualified judges assess anonymized texts under the same rubric, how reproducible are their ratings?

3. Behavioral evidence

Readers can report aesthetic appeal, surprise, emotional intensity and reread intention. Memory can be tested later through recognition or reproduction. Such measures provide evidence about what a text actually does to readers and complement textual analysis. Experiments have found measurable effects of rhyme and meter on aesthetic and emotional response [3] and of meter on recognition memory [4]. Newer work on English poetry also shows substantial individual differences in creativity judgments [6].

Historical influence belongs to a fourth, retrospective category. Citation, memorization, imitation, teaching, translation and long-term survival require time. The CAI can estimate anchor potential. Historical influence becomes evidence only after reception history exists.

The Poetic Excellence Index: 100 points

The PEI fixes ten dimensions before any poem is scored. Each applicable dimension receives 0–10 points and requires a textual reason. Two outputs are reported: Baseline PEI, with equal weighting across all applicable dimensions after the task profile is locked, and Task-adjusted PEI, which gives Central dimensions twice the weight of Supporting dimensions while excluding genuinely Inapplicable dimensions.

Before scoring: define the poem’s task profile

Record form or mode, scale, primary artistic task, secondary tasks when present, dominant mechanisms and comparator class. State the primary task as: This poem attempts to [effect] through [mechanisms] within [form/scale]. Then classify each PEI dimension as Central (2), Supporting (1) or Inapplicable (0).

Baseline PEI = 100 × Σ(active × score) / [10 × Σ(active)], where active = 1 for Central and Supporting dimensions. Task-adjusted PEI = 100 × Σ(relevance × score) / [10 × Σ(relevance)], where relevance = 2 for Central, 1 for Supporting and 0 for Inapplicable. A weak performance can never be declared inapplicable merely because it is weak. Genre creates expectations rather than commandments. The task profile is inferred from text, form, paratext and genre before authorial explanations are considered, and the profile must be locked before quality scores are assigned. In a formal study, separate raters should classify the profile and score execution whenever practical.

Equal weighting across applicable dimensions is the prespecified baseline for Method Version 1.1. Ten equal weights make the first version transparent and easy to falsify. Validation should compare Baseline PEI against Task-adjusted PEI on held-out poems, estimate criterion × profile interactions, measure agreement on task-profile classification, test redundancy and factor structure, and reject adaptive weighting if it fails to improve out-of-sample agreement.

Scoring precision. The rubric defines anchor regions rather than pretending that every integer has an independently validated meaning. Within a region, use the lower score when the stronger description is only partly satisfied and the higher score when it is consistently satisfied. Every one-point distinction requires a textual reason. A future validation study should preregister this baseline, sample sizes, exclusion rules, thresholds and analysis plan before data collection.

Criterion 0–3: failure 4–7: competent 8–9: major 10: extraordinary
1. Formal control Form repeatedly fights the poem: unstable lineation, accidental rhythm, unresolved constraint, or arbitrary layout. The chosen form works and is mostly controlled, with some passages that feel merely serviceable. Form creates pressure, expectation or contrast that materially strengthens meaning. Changing the governing formal decisions would damage several effects at once; control is sustained throughout.
2. Diction: precision, idiomatic force and purposeful deviation Cliché, padding, strained syntax, decorative vocabulary, or wording chosen mainly to satisfy form. Generally precise and idiomatic, with occasional generic or replaceable phrasing. Specific, economical diction with a distinctive voice and very little verbal slack. Word choice feels both surprising and inevitable; native expert readers find no forced or gratuitous phrase.
3. Sound and prosodic design Sound patterns are absent where the form demands them, mechanically repetitive, or audibly distort syntax and sense. Competent cadence, rhyme, stress, alliteration, assonance or free-verse rhythm; effects are local. Sound patterns organize attention and meaning across the poem. Phonetic and rhythmic design operates at several scales and remains inseparable from the poem’s argument or emotion.
4. Originality and surprise Central language, images or turns are predictable, derivative or cliché-driven. Some fresh choices occur within a familiar conceptual frame. The poem repeatedly produces earned surprise through language, thought or structure. Its major discoveries are difficult to paraphrase as borrowed formulas and remain convincing after surprise fades.
5. Semantic economy and density Many words do one job; paraphrase loses little. Several details carry more than one function, though stretches remain explanatory or redundant. Words, images and formal relations repeatedly carry multiple active functions. Every major unit earns its space and participates in multiple semantic, formal or structural relations; removing or replacing key material causes disproportionate loss.
6. Imagery and perceptual specificity Abstractions announce feelings or ideas without enough concrete pressure. Images are clear and relevant but sometimes familiar or illustrative. Specific images actively transform the poem’s argument or emotion. The image system is precise, memorable and generative: later meanings depend on earlier sensory details.
7. Figurative or conceptual coherence Figures, concepts or governing relations conflict accidentally, collapse under scrutiny, or require external rescue. The figurative or conceptual system is intelligible and mostly consistent, with limited development. Figures or concepts develop and interact coherently, surviving close testing while opening further consequences. The governing figurative or conceptual system sustains the poem’s full design with exceptional internal logic, generative power and textual accountability.
8. Structural design and consequence The governing structural logic is unclear, accidental or inert; arrangement contributes little to the poem’s task. A coherent structural principle is present and mostly functional, though some units feel replaceable or weakly consequential. Progression, recurrence, stasis, fragmentation, circularity or openness materially shapes the poem’s effects. The governing structural logic—progression, recurrence, stasis, fragmentation, circularity or openness—feels necessary; reordering, deleting or resolving a major unit damages several effects at once.
9. Interpretive yield and necessity Apparent depth depends mainly on vagueness, external biography or explanations unavailable in the text. The poem supports a stable reading with some meaningful payoff, though parts remain predictable, underdeveloped or weakly necessary. Interpretive payoff is substantial: several text-grounded readings may interact productively, or a deliberately clear meaning may acquire unusual force through exact formal and verbal relations. Interpretive yield is maximal for the poem’s artistic task: layered meanings remain tightly text-bound, or radical clarity achieves such necessity that further complication would weaken the work.
10. Intended affective or rhetorical force The intended affective or rhetorical effect is unclear, contradictory or weakly produced by the text. The intended effect is legible and partly achieved, with uneven force or limited textual support. The poem produces a strong, textually motivated affective or rhetorical effect appropriate to its artistic task. The intended effect is exceptionally precise, powerful and inseparable from the poem’s formal and verbal design.

These criteria are deliberately plural. Rhyme can raise a score when its execution contributes to sound, structure or memory; absence of rhyme carries no penalty. Length also carries no bonus. A compact epigram can score as highly as a long philosophical poem if it solves its own artistic problem with equal rigor.

The PEI works best after the kind of textual examination described in the complete close-reading method. Related concepts such as conceit, symbolism and metaphysical poetry help judges identify mechanisms while keeping the rubric device-neutral.

Provisional 99–100 disqualifiers

These ceiling rules remain provisional until blinded validation tests them. A near-perfect score should survive hostile close reading. For 99 or 100, the governing rule is no cheap point of attack.

  • A poem cannot receive 99+ if at least 25% of a prespecified native-expert panel independently identifies the same passage as forced or unnatural diction and supplies materially convergent textual reasons.
  • It cannot receive 99+ if the same threshold independently identifies a substantial line as existing chiefly to complete rhyme or meter.
  • It cannot receive 99+ if its central metaphor collapses under literal examination and the breakdown produces no compensating artistic effect.
  • It cannot receive 99+ if its supposed interpretive depth becomes available only after an external explanation by the author.
  • It cannot receive 99+ if at least 25% of the expert panel independently marks the same stanza as materially weaker than the rest for materially convergent reasons.

These are ceiling rules. The remaining criteria still determine the full score, while a conspicuous defect caps the result below 99.

Author-masked worked examples

The examples below come from public-domain English poetry. Author names are hidden at first, though familiar readers may recognize famous lines. These demonstrations inspect observable properties; they are not substitutes for a standardized blind experiment.

Example A

That time of year thou mayst in me behold
When yellow leaves, or none, or few, do hang
Upon those boughs which shake against the cold,
Bare ruined choirs, where late the sweet birds sang.

The phrase compresses season, architecture, religion, silence and remembered birdsong into one image. Its consonants and stresses make the emptiness audible. On PEI terms, the line is strong in semantic density, imagery, sound and figurative coherence.

Source: William Shakespeare, Sonnet 73.

Example B

Because I could not stop for Death –
He kindly stopped for me –
The Carriage held but just Ourselves –
And Immortality.

The conceptual surprise lies in courtesy: death becomes the party who observes social form. The syntax is plain, while the personification changes the emotional temperature of an enormous subject. The strength can be described through originality, compression, diction and tonal control before any biographical context is supplied.

Source: Emily Dickinson, “Because I could not stop for Death –”.

Example C

Death, be not proud, though some have called thee
Mighty and dreadful, for thou art not so;
For those whom thou think’st thou dost overthrow
Die not, poor Death, nor yet canst thou kill me.

The apostrophe creates an argument at once. Addressing Death directly converts metaphysical fear into rhetorical contest. The opening earns force through structure and stance before the poet’s name is known.

Source: John Donne, Holy Sonnet X.

Example D

Glory be to God for dappled things –
For skies of couple-colour as a brinded cow;
For rose-moles all in stipple upon trout that swim;
Fresh-firecoal chestnut-falls; finches’ wings;

The line begins with praise and immediately narrows into a precise visual principle: variegation. Its stresses and clustered consonants announce the sound world that the poem develops. Formal recognizability and phonetic design are observable before historical reputation enters.

Source: Gerard Manley Hopkins, “Pied Beauty”.

Full author-masked scoring example

The following demonstration uses a complete public-domain poem with the author hidden until the end. This is a single-rater illustration, not validation data.

Task profile: short symbolic lyric · Primary task: expose destructive hidden desire through a compressed symbolic image · Dominant mechanisms: image, metaphor, sound, compression and structural movement · Comparator class: short symbolic lyrics · Profile confidence: high.

O Rose thou art sick.
The invisible worm,
That flies in the night
In the howling storm:

Has found out thy bed
Of crimson joy:
And his dark secret love
Does thy life destroy.

PEI criterion Relevance Score Textual reason
Formal control Central 9 Two compact quatrains create a controlled movement from diagnosis to destruction.
Diction: precision, idiomatic force and purposeful deviation Central 9 Short, plain words carry the poem; little phrasing feels replaceable.
Sound and prosodic design Supporting 8 Worm/storm and joy/destroy bind the turns without making the diction feel mechanically driven.
Originality and surprise Central 9 The invisible worm and secret love fuse disease, desire and violation in an unexpected relation.
Semantic economy and density Central 10 Eight lines sustain botanical, erotic, bodily and moral readings with almost no exposition.
Imagery and perceptual specificity Central 9 Rose, worm, night, storm and crimson bed provide a concrete image system.
Figurative or conceptual coherence Central 9 The central image remains internally legible while supporting several text-grounded interpretations.
Structural design and consequence Central 9 The poem moves from sickness to agent, invasion, bed and final destruction.
Interpretive yield and necessity Central 9 Multiple readings arise from relations inside the poem rather than from biographical rescue.
Intended affective or rhetorical force Supporting 9 The compressed image sequence produces a controlled sense of menace and corruption appropriate to the poem’s task.
Results 10 active dimensions Baseline PEI: 90.0/100 Task-adjusted PEI: 90.6/100 · Profile confidence: high.

Source: William Blake, “The Sick Rose.”

A stronger validation uses matched texts by form, scale, mode and dominant artistic task whenever possible. Pairwise choice can then test whether the rubric agrees with independent preference without rewarding scale. Delayed recall, recognition and reread intention belong to the behavioral validation layer rather than criterion 10 itself.

The Canonical Anchor Index

PEI measures artistic execution. CAI measures prospective anchor potential. Separate scores keep celebrity out of artistic evaluation and keep a new poem from being penalized for its short reception history.

Each CAI dimension uses the same provisional 0–25 scale: 0–5 very weak, 6–10 limited, 11–15 substantial, 16–20 strong, 21–24 exceptional, 25 extraordinary ceiling. These bands are anchor regions. Within a band, every one-point distinction requires a concrete textual reason; the scale should not be read as empirically validated interval measurement. Scores above 20 require a concrete textual or reception-independent reason; 25 requires no obvious weak point within that dimension.

CAI dimension Question Maximum
Anchor strength Does the poem contain lines, images or formulations that remain intact and meaningful when recalled outside the full text? 25
Thematic reach Can the poem speak beyond its immediate occasion without dissolving into generic abstraction? 25
Formal recognizability Does the poem have a distinctive structure, voice or pattern that readers can identify and transmit? 25
Cross-context portability Can the poem sustain rereading across generations, situations, media or languages while retaining a coherent core? 25

A perfect epigram may earn PEI 96 and CAI 72; an equally accomplished poem with broader thematic reach and more portable anchor lines may earn PEI 96 and CAI 91. The second number describes a different kind of cultural affordance.

Historical influence is recorded separately. For a new work it should be marked “pending observation.” Decades of quotation, teaching, imitation, translation or adaptation may later supply evidence. That evidence belongs to reception history.

Interactive PEI and CAI scorecard

Methodological independence. Danil Rudoy developed this framework. Version 1.1 is fixed before any formal evaluation of Rudoy’s own poems. Future revisions receive a new version number and changelog entry; prior results remain archived. Any published comparison involving Rudoy’s work should rely on independent blinded raters, publish author self-scores separately, and release the protocol and anonymized aggregate data when available.

Version history. Version 1.1.1 (September 2026) is an implementation-conformance patch aligning the interactive scorecard with the published Version 1.1 formulas and lock rules. Version 1.1 added locked task profiles, applicable-dimension Baseline PEI, Task-adjusted PEI, CAI anchors, operational panel thresholds and explicit scoring-precision rules. Version 1.0 introduced the ten-dimension baseline framework.

Provisional PEI interpretation

PEI Provisional interpretation
99–100 Extreme ceiling range
97–98 Extraordinary
95–96 Exceptional
90–94 Major
80–89 Strong
Below 80 Requires empirical calibration before finer labels are claimed

Use the scorecard after close reading. A personal score is exploratory. Research-grade claims require independent blinded raters and reported agreement.

Baseline PEI: Unrated Task-adjusted PEI: Unrated Active dimensions: Unrated Profile confidence: Unrated CAI: Unrated

Poem Task Profile

Unlocked

Poetic Excellence Index

Lock the task profile first. Relevance: Central = 2, Supporting = 1, N/A = 0. A weak criterion cannot be made N/A merely because it scores poorly. Every score requires textual evidence.

Canonical Anchor Index

CAI anchors: 0–5 very weak, 6–10 limited, 11–15 substantial, 16–20 strong, 21–24 exceptional, 25 extraordinary ceiling.

What the score can say

The PEI is a structured argument about artistic performance. Two readers may agree that a poem displays exact formal control, novel imagery and remarkable compression while still preferring different works.

Future fame belongs to reception history. Canon formation depends on institutions, transmission, education, politics, translation, chance and historical need. PEI therefore scores intrinsic poetic execution without a celebrity or fame component.

Empirical findings about one device require bounded interpretation. Obermeier and colleagues found that rhyme and regular meter affected aesthetic and emotional ratings in their experimental materials [3]; Van Peer found an effect of meter on aesthetic pleasure and recognition memory [4]. These results justify measuring formal patterning and reader response. Great poetry may use regular meter, free verse, rhyme, partial rhyme, or other organizing principles.

Research on parallelism and deviation shows that recurring patterns and expectation violations can be quantified and can predict aspects of comprehension and aesthetic evaluation [5]. Studies of English poetry find that aesthetic appeal, surprise and individual reader differences contribute to creativity judgments [6] [7]. Expert–novice eye-tracking research also shows that expertise changes poetry-reading behavior [8]. Together, these findings support a program of measurement with explicit uncertainty.

The central objection therefore has a precise answer. Two people are free to prefer different poems. Preference can remain personal. Yet the ability of a text to sustain its form, create nontrivial relations, produce meaning economically, remain memorable and receive reproducibly strong judgments from independent qualified readers consists of observable facts. Subjectivity creates variance; it does not erase signal.

A hostile critic can still say, “I still don’t like this poem.” That statement records preference. A substantive challenge must identify an undefined criterion, circular evidence, cherry-picked texts, unblinded raters, contamination by author fame, or invented data. Any such failure requires revision before the method deserves confidence.

Frequently asked questions

Is poetry subjective?

Poetry includes subjective response, and individual differences matter. It also contains observable textual features and produces responses that can be measured across readers. The useful question is how much reliable signal remains after variance is measured.

Can the PEI prove that a poem is great?

PEI is a falsifiable rubric for making evaluative claims explicit and testable. Its credibility depends on blind application, independent raters, reliability estimates, matched comparisons and replication.

Does a poem need rhyme or meter to score highly?

Rhyme and meter are possible formal resources. Free verse can demonstrate equally strong formal control through cadence, syntax, lineation, recurrence and other constraints. A device earns credit for what it accomplishes in the particular poem.

Can a new poem receive a high Canonical Anchor Index?

It can receive a high prospective CAI based on anchor strength, thematic reach, formal recognizability and cross-context portability. Historical influence remains unscored until evidence accumulates over time.

Why use blind raters?

Blind presentation reduces prestige, biography and expectation effects. If a method is meant to judge the text, the first pass should conceal information that can bias judgments before the text is read.

References

  1. Amabile, Teresa M. “Social Psychology of Creativity: A Consensual Assessment Technique.” Journal of Personality and Social Psychology 43.5 (1982): 997–1013. APA PsycNet.
  2. Kaufman, James C., John Baer, and Jason C. Cole. “Expertise, Domains, and the Consensual Assessment Technique.” The Journal of Creative Behavior 43.4: 223–233. Wiley Online Library.
  3. Obermeier, Christian, et al. “Aesthetic and Emotional Effects of Meter and Rhyme in Poetry.” Frontiers in Psychology 4 (2013): 10. Frontiers.
  4. Van Peer, Willie. “The Measurement of Metre: Its Cognitive and Affective Functions.” Poetics 19.3 (1990): 259–275. ScienceDirect.
  5. Menninghaus, Winfried, et al. “Parallelisms and Deviations: Two Fundamentals of an Aesthetics of Poetic Diction.” Philosophical Transactions of the Royal Society B 379 (2024): 20220424. PubMed / U.S. National Library of Medicine.
  6. Chaudhuri, Soma, Alan Pickering, Maura Dooley, and Joydeep Bhattacharya. “Beyond the Words: Exploring Individual Differences in the Evaluation of Poetic Creativity.” PLOS ONE 19.10 (2024): e0307298. PLOS ONE.
  7. Chaudhuri, Soma, Alan Pickering, and Joydeep Bhattacharya. “Evaluating Poetry: Navigating the Divide between Aesthetical and Creativity Judgments.” The Journal of Creative Behavior 59.1. Wiley Online Library.
  8. Fokin, Danil, Stefan Blohm, and Elena Riekhakaynen. “Reading Russian Poetry: An Expert–Novice Study.” Journal of Eye Movement Research 13.3. University of Bern.