What this method produces
This method codes arts into the Open Martial Arts Database and fights into the Open Combat Database. See where they meet: the fit.
The Eskirmological Analysis Protocol
A reproducible method for the behavioural classification of martial arts.
A procedure for turning a documented martial art into a quantitative, verifiable behavioural profile. It takes a defined body of source material as input and returns a map: a distribution of that art’s behaviour across a fixed coordinate system, together with the reliability statistics needed to judge how much trust the map deserves.
Purpose and scope
The Eskirmological Analysis Protocol (EAP) is a procedure for converting a documented martial art into a quantitative, verifiable behavioural profile. It takes a defined body of source material (a corpus) as input and returns a map: a distribution of that art’s behaviour across a fixed coordinate system, together with the reliability statistics needed to judge how much trust the map deserves.
The protocol has one governing ambition. Two competent analysts, working independently on opposite sides of the world from the same corpus and the same version of this manual, should recover statistically equivalent maps. Where they do not, the disagreement should be measurable, locatable, and correctable. This is the replication contract on which everything else rests.
The protocol is deliberately silent on nationality, lineage, terminology, philosophy, mythology, intended application, and instructor interpretation. It measures only observable action. In this it follows the ethological convention of describing what an organism does before asking what the behaviour is for (Tinbergen, 1963; Lehner, 1996), and the content-analytic convention of coding manifest content against a manual fixed in advance of coding (Krippendorff, 2018; Neuendorf, 2017).
Theoretical position
The EAP sits at the intersection of five established traditions, none of which is usually applied to martial arts. Each contributes a piece of methodological machinery that the protocol reuses rather than reinvents.
The payoff of assembling these is that Eskirmology stops being a taxonomy of styles and becomes a coordinate system. Taxonomies sort things into boxes. A coordinate system lets you measure distances, identify neighbours, quantify gaps, and test claims. Statements such as “Wing Chun and Taijiquan are closer than practitioners assume” or “Karate and Taekwondo occupy nearly the same coordinates” cease to be opinion and become measurable quantities.
The coordinate system
Every combat behaviour is described by a value on each of three orthogonal facets. The facets answer three questions that are independent of one another: an answer to one places no constraint on the answers to the others.
| Facet | Question it answers | Governs |
|---|---|---|
| T Temporal | How is initiative acquired? | The polarity of the exchange |
| S Spatial | How is distance managed? | The range at which the problem is solved |
| M Modal | By what mechanism is the opponent affected? | The physical means of resolution |
T is not a count of how many people are involved. It describes the temporal intention of the system: who authors the problem being solved.
Unilateral
The system imposes its own solution. Initiative-first. The practitioner creates the problem the opponent must answer.
Bilateral
The system resolves the opponent’s action. The incoming action is accepted, redirected, borrowed, or exploited before the system commits to its own resolution.
This is the original polarity of the 2008 formulation, and it is the facet with the highest coding risk, because whether a given action initiates or responds often depends on the exchange it sits inside rather than on the action itself.
Hit and away
The system operates outside sustained contact. It strikes and recovers distance; contact is momentary.
Stand and cope
The system operates at established contact. It holds position and manages the exchange in place, at bridge, clinch, or trapping range.
Close and displace
The system operates at body attachment. It closes the remaining gap and displaces the opponent’s structure.
M is a two-level hierarchy in which manipulation and percussion each split into a pair. The hierarchy earns its place operationally: an analyst who cannot yet decide between M3 and M4 can still reliably record “manipulation, non-percussion,” and the finer distinction becomes a second, separable coding decision. Agreement can then be reported at both levels.
Hands
Upper-limb percussive delivery.
Feet
Lower-limb percussive delivery.
Small body manipulation (SBM)
Control exerted on local structures: joints, wrists, the opponent’s frame at close range; redirection and leverage.
Great body manipulation (GBM)
Whole-body displacement: throwing, projection, sweeping, off-balancing to the ground.
The facets combine into a grid of 2 (T) × 3 (S) × 4 (M) = 24 atomic coordinates. This grid is the Zwicky box for combat behaviour: the complete space of behavioural addresses that any technique can occupy. A coordinate is written as a triple, for example (T1, S1, M1), an initiative-first, out-of-contact hand strike (a lead jab thrown as an opening).
The 24 cells are enumerated a priori. Which cells a given art actually occupies, which it neglects, and which appear to be empty across all arts, are empirical findings produced by the protocol, not assumptions built into it. Zwicky’s discipline applies: enumerate everything first, then look.
The ontology, and reconciliation with earlier drafts
The definitions in Section 3 constitute the authoritative ontology for Version 1.0. Earlier working notes used a flat, four-value M facet labelled percussion / leverage / redirection / projection. Version 1.0 supersedes that labelling with the percussion–manipulation hierarchy for two reasons: it distinguishes hand from foot percussion, which the flat scheme could not, and it groups leverage and redirection under a single manipulation class.
For anyone migrating older analyses, note that migration cannot be automatic. Percussion primitives coded under the old scheme must be re-examined to separate M1 from M2, and the old leverage and redirection categories collapse into M3 while projection maps to M4. In practice, treat pre-hierarchy data as a different manual version and re-code rather than translate, to avoid silent cross-version contamination.
An ontology in this sense is a shared, explicit specification of the concepts and their relations, so that different people mean the same thing by the same code (Gruber, 1993). The coding manual is where that specification is made operational.
Units of analysis
A recurring error in behavioural coding is to force every facet onto the same unit. The EAP uses two, matched to what each facet actually measures.
The unit for S and M
The smallest meaningful combat action that still carries a spatial and a mechanical identity: rotate forearm to intercept, reap supporting leg, drive rear hand along centreline. S and M are close to intrinsic to the primitive.
The unit for T
A minimal initiate-and-respond cycle. Initiative is a relational property of the exchange, so it is read from the sequence, not from an isolated action.
Each coded record carries a full (T, S, M) triple, which places it at one of the 24 coordinates. The primitive supplies S and M; the exchange it belongs to supplies T.
The same jab can be an opening (T1) or a counter (T2). If T is coded from the isolated primitive, two analysts will disagree, and agreement on T will be lower than on S or M. The protocol offers two treatments; a corpus must declare which it uses.
Whichever is used, the choice is part of the corpus manifest and the manual version, so replication is exact.
The pipeline
The protocol is a fixed sequence from corpus to map, tightening the original nine-stage sequence into twelve reproducible steps.
Register the corpus
Produce the manifest: exact source, edition, syllabus version, and precise statement of what is included and excluded.
Segment
Divide the corpus into exchanges (for T) and, within them, into primitives (for S and M).
Decompose
Reduce each documented technique to its constituent primitives. Age uke becomes raise arm, rotate forearm, intercept, establish contact. O soto gari becomes enter, off-balance, leg reap, projection. Nothing larger than a primitive is coded.
Code
Assign each unit its (T, S, M) coordinate using the decision trees in Section 7. Coders work independently and do not confer.
Check reliability
At least two coders independently code a shared sample. Compute agreement per facet. If any facet falls below threshold, revise the manual, re-train, and re-code. Do not proceed on an unreliable facet.
Count
Tabulate frequencies for each coordinate and each facet value.
Distribute
Convert counts to proportions, with interval estimates.
Build the signature
Assemble the distribution over the 24 coordinates and its three facet marginals. This is the map.
Compress the formula
Derive the human-readable descriptor by the deterministic rule in Section 9.3. The formula is derived, never assigned.
Analyse coverage
Map the signature onto the 24-cell grid; identify occupied, sparse, and empty cells.
Report confidence
State sample sizes, agreement coefficients, and interval widths.
Deposit
Lodge the replication packet so the result can be reproduced and challenged.
The coding manual
The single most important document in the protocol. The formulae are downstream of it. For every facet there is an objective decision tree, applied without reference to what the technique is called or for.
Reliability and validity
Because the facets are independent and are coded by different decisions, agreement must be reported separately for each facet (αT, αS, αM), and, for M, optionally at both the class level (percussion vs manipulation) and the leaf level (M1–M4). A single pooled figure would hide the facet that is failing, which is usually T.
Report the coefficient, the number of coders, and the size of the reliability sample.
Adopt the Landis and Koch (1977) benchmarks as the community convention, while treating them as guides rather than law.
| Coefficient | Interpretation |
|---|---|
| < 0.00 | poor |
| 0.00–0.20 | slight |
| 0.21–0.40 | fair |
| 0.41–0.60 | moderate |
| 0.61–0.80 | substantial |
| 0.81–1.00 | almost perfect |
Reliability is agreement between coders. Validity is whether the codes measure what they claim to. High agreement on a badly designed manual is reliably wrong. Two safeguards:
From counts to signature
For a facet value observed k times in n coded units, the point estimate is p̂ = k / n. Report an interval around it. Use the Wilson score interval (Wilson, 1927) rather than the textbook normal-approximation interval, because proportions in this work are frequently near 0 or 1 (an art may make almost no use of M4), and the normal approximation misbehaves badly at the extremes. State the confidence level (95% by default) alongside every proportion.
The facet distributions are compositional: they are parts of a whole that sum to one. This is not a technicality. Ordinary Euclidean distance and ordinary correlation give misleading answers on compositional data, a hazard documented at length by Aitchison (1986). For any distance or comparison between signatures, use a method appropriate to compositions: a log-ratio transformation in the Aitchison tradition, or the information-theoretic divergence in Section 10.1. Recording this in the manual protects the community from a whole class of silent statistical error.
The full signature is the distribution over the 24 coordinates. The three facet marginals are summaries of it. Capturing the complete (T, S, M) triple per unit is what makes the joint distribution, and therefore genuine gap analysis, possible; coding the facets separately would discard the interaction structure and reduce the map to three disconnected histograms.
The formula must be reproducible, so it is generated by a rule, not by judgement.
Fix a threshold τ as part of the manual version (default τ = 0.15). Within each facet, list every value whose proportion is ≥ τ, in descending order of proportion, joined by “/”. Concatenate the three facets in T, S, M order.
Suppose an art’s marginals are T2 = 0.84; S2 = 0.71, S1 = 0.20, S3 = 0.09; M1 = 0.44, M3 = 0.38, M2 = 0.10, M4 = 0.08. With τ = 0.15 the formula is [T2][S2/S1][M1/M3]. Change τ and the formula may change, which is why τ travels with the manual version and the raw signature is always deposited so any reader can recompute at a different threshold.
Comparison, coverage, and gap analysis
The practical payoff: profiling a skillset against the whole solution space and finding what is missing.
Because signatures are probability distributions, use a proper divergence. The recommended default is the Jensen–Shannon divergence (Lin, 1991): it is symmetric, always defined (it tolerates zero cells, which Kullback–Leibler does not), and bounded in [0, 1] when computed with base-2 logarithms, so distances are comparable across pairs. Compute it on the 24-cell joint distribution for the fullest comparison, or per facet when you want to say precisely where two arts differ.
This turns qualitative claims into measured quantities. “Wing Chun and Taijiquan are closer than people assume” becomes a small JSD. “Judo and freestyle wrestling are architecturally near-identical, differing mainly in rule sets” becomes a JSD near zero on the T/S/M map even where their rule sets diverge. “Jeet Kune Do is an attempt to move freely between coordinates” becomes a signature that is unusually flat across the grid rather than concentrated, which the protocol can detect and report as low concentration (for instance a high entropy of the 24-cell distribution).
Lay the signature over the 24-cell grid and classify each cell:
The gap profile is the set of sparse and empty cells. For an individual practitioner, coding their personal repertoire the same way produces a personal coverage map, and the empty cells are development targets stated in behavioural terms rather than by style name. A practitioner can see, against the total solution space, exactly which regions of combat behaviour their training does not reach.
An art B complements art A when B’s mass concentrates in A’s gaps. Make this precise:
Let gA be A’s gap set (its sparse and empty cells). Rank candidate arts by the share of their signature that falls inside gA. The highest-ranked arts are A’s strongest complements.
This regenerates, as a computation, the complementary-art table: Shotokan [T1][S1][M1/M2], whose gaps are the bilateral and manipulation regions, is complemented by Aikido, Wing Chun, Taijiquan, and Hapkido, precisely the arts whose mass sits in T2 and M3/M4. Boxing [T1][S1][M1], empty across feet, clinch, and manipulation, is complemented by Muay Thai, Judo, and BJJ. The table stops being a matter of taste and becomes the output of a defined procedure on the maps.
Reproducibility apparatus
Given the same corpus, the same coding-manual version, and the same pipeline parameters (unit definitions, τ, coefficient choice), any competent coder pool should recover the same map within stated tolerances: per-facet agreement at or above the published threshold, and signature distances within the reported confidence intervals.
Every analysis is judged against this contract. A result that cannot be reproduced under it is not yet a result.
A registered corpus records, at minimum: the art and sub-style; the exact source and edition (syllabus document, named form, grading manual, filmed reference, with dates); the syllabus version; the inclusion and exclusion rules (which grades, which forms, which drills, and what was left out and why); the segmentation choices; and the T treatment (exchange-level or canonical-context). The test is simple: could a stranger obtain and analyse exactly the same material?
Version the coding manual with three-part semantic versioning, MAJOR.MINOR.PATCH:
Every map cites the manual version that produced it. Maps from different MAJOR versions are not directly comparable and must be re-coded before comparison.
The deposited unit of work. It contains: the corpus manifest; the manual version; the full coded dataset (every unit with its (T, S, M) triple and coder ID); the reliability sample and computed coefficients; the derived signature, formula, and gap profile; and all parameters (τ, confidence level, coefficient). A schematic record for one coded unit:
{ "corpus_id": "shotokan-jka-heian-2026-01", "manual_version": "1.0.0", "unit_id": "heian-nidan-ex-014-p2", "unit_type": "primitive", "exchange_id": "heian-nidan-ex-014", "T": "T1", "S": "S1", "M": "M1", "coder_id": "A", "notes": "reverse punch as opening; hand percussion at range" }
Machine-readable records make the map regenerable by anyone: the signature is a deterministic function of the coded dataset plus parameters.
As packets accumulate, they form a public Eskirmology Corpus: a database of coded analyses, each carrying its manifest, coding decisions, sample size, signature, confidence metrics, and manual version. This follows the transparency-and-openness conventions now standard in reproducible research (Nosek et al., 2015): shared data, shared instruments, shared analysis. The gold-standard validation is the multi-analyst study, for example four analysts in four countries coding the same 26 Shotokan kata against the published manual. If their signatures converge and leaf-level agreement clears the threshold, the taxonomy has been shown not to depend on a single expert’s reading. At that point Eskirmology is a measurement protocol, not an interpretation.
The confidence report
Every published map carries a short, fixed report so a reader can judge it at a glance.
n = 2,438 primitives / 611 exchanges; αT = 0.79, αS = 0.93, αM = 0.91 (three coders, reliability sample 300); widest 95% Wilson interval ±4.1%; corpus shotokan-jka-heian-2026-01; manual 1.0.0. A reader learns immediately that S and M are trustworthy, that T is close to threshold and should be read with mild caution, and that the map is reasonably sharp.
Known limitations and open problems
Stated plainly, because a protocol that hides its weaknesses cannot be trusted with its strengths.
These are the productive edges of the method. Each is a place where the shared Corpus, over time, can test and refine the instrument.
Glossary
- Atomic decomposition
- Reduction of a documented technique to its smallest coded actions (primitives).
- Corpus
- The defined, documented body of source material analysed for one art.
- Coordinate
- A single (T, S, M) triple; one of the 24 cells of the solution space.
- Ethogram
- A documented inventory of an organism’s discrete behaviours; here, of a fighting system.
- Formula
- The compressed, human-readable descriptor derived from a signature by the τ rule.
- Gap profile
- The sparse and empty cells of a signature relative to the solution space.
- Map / signature
- The distribution of an art’s coded behaviour over the 24 coordinates, with its facet marginals.
- Primitive
- The smallest meaningful combat action carrying a spatial and mechanical identity; the unit for S and M.
- Exchange
- A minimal initiate-and-respond cycle; the unit for T.
- Replication packet
- The complete, deposited record that lets anyone regenerate and challenge a map.
- Solution space
- The full 2 × 3 × 4 grid of behavioural coordinates.
From protocol to coded corpus
The instrument is defined. The demonstration is a full-system case study coded against this manual, deposited as a replication packet for others to challenge.