ARKONE
A Route Model, Asked Blind: AI Route Selection for Regional Indian Aviation

A Route Model, Asked Blind: AI Route Selection for Regional Indian Aviation

August 17, 2026 · 69 min read

S
Sobin George Thomas

We built a route-selection model for regional Indian aviation from one public dataset. Asked blind, it placed 8 of a carrier's 10 routes in its top 100 of 2,743 city pairs.

Ask a model which routes a turboprop airline should fly. Give it nothing except a public spreadsheet. Before it earns the right to suggest anything new, it has to pass one test: it must rediscover, on its own, the routes a veteran planning team already chose.

In August 2026 we ran that test against a regional Indian carrier we admire. The model read one public source, scored 2,743 city pairs, and placed eight of the carrier’s ten routes in its top 100. It had been told nothing about the carrier. Then it named three city pairs the carrier does not fly, with the arithmetic behind each score laid open.

2,743
City pairs scored
Every pair inside the turboprop envelope, nationwide
8
Carrier routes in the blind top 100
Of the ten routes it flies; chance expects 0.4
0
Carrier inputs given to the model
Public DGCA data only

This report documents the whole build: the market and its operators, the four questions the model asks, the unglamorous data work that made a government file usable, the blind validation, the findings, a national atlas of where the thin routes live, and the limits we state before anyone asks. A note on provenance before anything else. This is an independent analysis. No carrier commissioned it, shared data for it, or had any involvement in it; the study claims no relationship with the operator it models, and every number in it comes from public data. The work was done to prove a method, on our own time.


The market: a decade of thin routes

Indian domestic aviation carried 166.9 million passengers in 2025, by the DGCA’s own city-pair records. The headline number hides its own structure. A large share of those passengers move between a dozen metro airports on 180-seat narrowbody jets. The rest move on everything else: the medium cities, the pilgrimage towns, the island strips, the state capitals that got an airport before they got demand.

That second tier is where regional aviation lives, and it behaves differently. A Mumbai–Delhi planner asks how many frequencies the slots allow. A regional planner asks a harder question: does this pair produce enough passengers to fill a small aircraft daily, without producing so many that a low-cost jet operator walks in and takes the market? The band between those two failure modes is narrow. Most city pairs in India sit outside it, on one side or the other.

The policy environment has decided to widen that tier. In 2026 the Union Cabinet approved the Regional Connectivity Scheme, Modified UDAN, with a ₹28,840 crore outlay running from FY 2026-27 to FY 2035-36. The scheme funds new airports out of unserved airstrips and carries viability gap funding for the operators willing to serve them. The first UDAN decade built the runway for the second: by mid-2026 the DGCA’s own daily dashboard counts 657 RCS routes awarded across 93 operationalised airports, flown by 10 operators. Whatever one thinks of subsidy economics, the direction is unambiguous: more airports, more thin pairs, more decisions about which of them deserve an aircraft.

The regional connectivity build-out
657
UDAN routes awarded and active
DGCA daily statistics, July 2026
93
RCS airports operationalised
Modified UDAN targets 100 more by 2036
₹28,840 cr
Modified UDAN outlay
FY 2026-27 to FY 2035-36, Union Cabinet

More decisions is the operative phrase. Every new airport multiplies the combinations a network planner has to consider. Ten cities make 45 possible pairs. Fifty cities make 1,225. A carrier that intends to grow from a dozen cities to fifty chooses from a combinatorial field that grows quadratically, while the planning team stays the same size and the shortlist stays a comforting fiction.

And the choice is expensive to get wrong in both directions. Launch a pair that cannot fill the aircraft, and the route bleeds cash for the two or three seasons it takes to admit the mistake. Skip a pair that was quietly maturing, and a competitor’s schedule announcement converts your hesitation into their base. Route selection is the highest-stakes recurring decision a regional carrier makes, and it recurs faster as the network grows.

The aircraft shapes the problem as much as the market does. A 70-seat turboprop exists because of a stubborn piece of economics: below a certain demand density, a jet’s seat costs cannot be covered at fares a regional market will pay, while a turboprop’s lower trip cost and slower cruise fit sectors under two hours almost exactly. The price of that fit is a narrow operating envelope. Fly the turboprop too short and taxi-climb-descend eats the block time; fly it too long and the cruise-speed disadvantage compounds against jet competition on the same pair. Route selection for a turboprop operator is therefore a bounded search, and bounded searches are precisely where scoring models earn their keep: the envelope is definable, the candidates are enumerable, and the trade-offs repeat.

The conventional answers to this problem come in two shapes. The first is experience: a planning head who has watched Indian city pairs for twenty years and carries the map in their head. This works, and the carriers that have it should treasure it. It does not scale past the one person, it does not transfer when they leave, and it cannot audit itself. The second shape is the consultancy market-study, produced over eight weeks, delivered as a slide deck, stale within two quarters, and structurally incapable of saying how it reached its ranking.

We wanted to demonstrate a third shape. A small, legible model that reads public data, scores every feasible pair, shows its arithmetic, and can be re-run every month for the cost of a coffee. Not as a replacement for the twenty-year planning head. As the instrument panel in front of them.


The operators: who actually flies regional India

Before scoring routes, score the field. The same DGCA publication set that records city-pair traffic also records airline-level operations, and reading it settles a question that sector commentary usually hand-waves: who is actually in the regional business, at what scale, and in which direction each is moving.

The dividing line is average stage length, the distance a carrier’s typical passenger flies, computed here as passenger-kilometres over passengers carried in the twelve months to May 2026. The mainline operators cluster between 950 and 1,200 kilometres. The regional operators cluster between 280 and 580. There is almost nothing in between: India’s domestic market is two markets wearing one licence category.

Operator Passengers, 12m to May 2026 Year on year Avg stage
IndiGo 107.5M +1% 948 km
Air India 25.5M +7% 1,076 km
Air India Express 19.3M +48% 1,003 km
Akasa Air 8.8M +12% 1,195 km
SpiceJet 5.1M +5% 1,012 km
Star Air 933,712 +33% 582 km
Alliance Air 587,147 −49% 421 km
FLY91 390,708 +94% 442 km
IndiaOne Air 27,928 −8% 320 km
flybig 3,989 −77% 285 km

Airline-level operational data lags the city-pair record by about two months in the public mirror, so this table’s window is the twelve months to May 2026; the route model in the rest of this study runs to July.

The regional tier deserves its own reading, because its five members are moving in five different directions.

The regional tier, trailing 12 months to May 2026 (bar = share of the tier's passengers)
Star Air · 933,712 pax · +33%
48%
Alliance Air · 587,147 pax · −49%
30%
FLY91 · 390,708 pax · +94%
20%
IndiaOne Air · 27,928 pax · −8%
2%
flybig · 3,989 pax · −77%
1%
Computed from DGCA monthly airline statistics via the india-aviation-traffic mirror

Star Air is the current volume leader, an ERJ operator out of Karnataka growing a third year on year at the longest stage length of the tier, the profile of a regional edging toward mainline thinness limits. Alliance Air, the state-owned incumbent that carried the UDAN mission for years, halved in a year; a fleet and financing story working itself out in the traffic numbers. FLY91 nearly doubled, the fastest growth in the tier, from the smallest ATR base. IndiaOne Air and flybig illustrate the tier’s mortality: single-digit-aircraft operations on the thinnest routes, one drifting, one effectively exiting. Five operators, five trajectories, one conclusion for a route model: the regional tier is an open competition, and the demand pool it contests is one the mainline carriers structurally cannot serve.

That churn is the strategic backdrop for everything that follows. Routes vacated by a shrinking operator become someone else’s opportunity within two seasons. A monthly-refreshed instrument sees the vacancy the month the capacity leaves.


The geography of thin routes: a national atlas

The model built for this study is deliberately carrier-shaped: its fourth question scores proximity to one carrier’s bases. Strip that question out, renormalise the other three, and the same scoring runs carrier-agnostically over the whole country. The result is an atlas: every city pair inside the turboprop envelope, scored on size, growth, and range alone, then grouped into six analytic zones by coordinate. The zones are bands of geography, not administrative regions, and pairs spanning zones are assigned by midpoint.

Zone Pairs in envelope In the 70-seat band Passengers, 12m Year on year
South 714 70 34.4M −2%
North 822 37 21.7M +2%
West 376 30 13.9M +3%
Northeast 182 19 4.9M −5%
East 142 15 2.1M −18%
Central 428 13 1.6M −9%

Two structural facts fall out of the table. The South is the deep end. It holds 70 of the country’s 184 in-band pairs, twice its nearest rival, on the largest passenger base. Peninsular geography does this: a lattice of mid-sized cities 250 to 700 kilometres apart, exactly the turboprop’s stride, with rail connections slow enough to lose the time comparison. The North is wide but shallow. More pairs in the envelope than any zone, but only 37 in the band, because northern demand concentrates on trunk routes into Delhi that outgrew 70 seats long ago, while the remainder is thinner than a daily rotation can justify.

The Northeast tells a third story worth a report of its own: 19 in-band pairs on terrain where flying replaces a day of driving, the strongest structural case for turboprops in the country, and the zone where operator mortality has been highest. The East’s −18% is the sector’s sharpest regional decline, concentrated in pairs that lost service rather than pairs that lost demand.

The atlas is deliberately published at zone level rather than as a national ranked list. A carrier-agnostic score is a screening lens, and the pairs at the top of each zone change meaning entirely depending on whose aircraft, whose bases, and whose competition you ask about. The instrument exists to answer that question per operator; the atlas exists to show where the questions are worth asking.


The carrier, and the question we set ourselves

The carrier in this study operates a fleet of six ATR 72-600 turboprops from two bases, one on the Konkan coast and one in the Deccan. It serves a network of around a dozen cities across western and southern India, a mix of metro connections, medium cities, and one island strip. A portion of its flying operates under UDAN route awards. Its stated ambition, in public interviews, is a fleet an order of magnitude larger, serving several times as many cities.

We chose this carrier for three reasons. The network is recent, which means every route in it was chosen deliberately by people who are still there, against current market conditions, with modern data. The fleet is uniform, six examples of one aircraft type, which collapses the modeling problem to a single operating envelope. The ambition is public, which means the question our model answers, which pair next, is a question this carrier will face dozens of times in the next five years.

There was also a fourth reason, and it is the honest one. We build AI systems for organisations, and the hardest part of that trade is the first conversation. Every vendor claims capability. Decks are free. We wanted to walk into aviation with something that runs, on data nobody could accuse us of massaging, validated against decisions we could not have known in advance. The way to earn a network planner’s attention is to show them a model that independently agrees with the calls they got right.

So we set the constraint that gives this study its title. The model would receive no information about the carrier at all. Not its route map, not its schedules, not its load factors, not its finances. One input only: the public record of how many people flew between every pair of Indian cities, every month, published by the Directorate General of Civil Aviation. If the model could not find the carrier’s network inside that file, blind, the model was wrong and we would not show it to anyone.

That constraint sounds like a handicap. It is actually the entire methodological point, and it is worth pausing on why.

Any model fitted to a carrier’s own data learns the carrier’s decisions, including its mistakes. It becomes an echo. A model built only on the public demand record has no access to the decisions, so when it reproduces them anyway, something real has been demonstrated: the decisions are recoverable from the demand data alone. The validation cuts both ways, and the second direction is the valuable one. It validates the carrier. A planning team that chose, from private judgement, the same routes a blind scorer finds in the public record has had its judgement independently confirmed by arithmetic. Where the scorer and the team disagree, the disagreement is precise, inspectable, and worth a meeting.

The blindness also settles the credibility question that kills most analytical vendor pitches. A ranking produced after studying the client’s network can always be accused, politely, of being reverse-engineered. A ranking produced before any contact, from data anyone can download, and shown with its weights exposed, cannot. It is falsifiable in the plain sense: anyone with a laptop can re-run it and check.

One scoping decision completes the setup. We modeled demand, deliberately, and nothing else. Fares, competitor schedules, slot availability, and crew logistics are all real, and all absent, and the limits section near the end of this study treats that absence at length. The short version: the public record supports a demand model honestly. Pretending it supports a profitability model would be the kind of overreach the aviation industry has learned to smell.


One data source, four questions

The model reads a single file: the DGCA’s monthly record of domestic passengers carried between every pair of Indian cities, from April 2015 to July 2026. We take the machine-readable mirror maintained in the open-source india-aviation-traffic repository, which aggregates the DGCA’s published monthly statistics into one continuous CSV. Eleven years of it: 66,453 monthly records across more than 1,800 distinct city pairs that have seen at least one scheduled passenger.

From that file, the model asks four questions about every city pair an ATR 72-600 can plausibly serve. Each question produces a sub-score between zero and one. A weighted sum produces a score out of 100, and every pair is ranked against every other. That is the entire model. There is no machine learning in it, no embedding, no black box. This was a design decision, and the reasoning behind each question carries most of the value, so here they are in full.

The four questions, weighted
Is the pair the right size for a 70-seater?34%
2,000 to 25,000 passengers a month. Enough to fill a small aircraft daily, too thin for a 180-seat jet to defend.
Is it growing?26%
The trailing twelve months against the twelve before them.
Can the aircraft do it comfortably?22%
250 to 700 km is the comfortable band. 150 and 900 are the hard limits.
Does it touch a base?18%
Where the aircraft, the crew and the engineers already sleep.

Question one: is the pair the right size? (34%)

A 70-seat turboprop flying a pair daily in both directions, at a healthy load factor, carries roughly 3,150 passengers a month. The model therefore rewards pairs moving between 2,000 and 25,000 passengers a month. Below 2,000, even a daily small aircraft flies with empty seats, and the sub-score falls away in proportion. Above 25,000, the pair can fill 180-seat jets, which means the low-cost majors either serve it already or will the moment anyone proves it, and the sub-score decays accordingly.

This band is the strategic heart of the model, which is why it carries the largest weight. Regional aviation is the business of pairs too big to ignore and too small to defend for a mainline jet operator. Everything else in the model refines this one idea.

Question two: is it growing? (26%)

The model compares the trailing twelve months, June 2025 through May 2026, against the twelve months before that. Flat demand scores 0.5. Growth of fifty percent or more scores the full 1.0, and decline falls away below the midpoint. Where the prior-year base is under 500 passengers, the model refuses to compute a growth rate at all and assigns the neutral 0.5, because percentage growth on a three-digit base is noise wearing a suit.

The intent is to catch waves rather than ponds. Two pairs may both move 8,000 passengers a month, but the one that moved 5,500 a year ago is telling you something the static one is not: demand on this pair is being discovered, and the discovery is not finished.

Question three: can the aircraft do it comfortably? (22%)

The ATR 72-600 is a sprinter. Its economics are at their best on sectors of 250 to 700 kilometres, roughly forty minutes to ninety minutes in the air. The model treats that band as ideal, scoring 1.0. Between 150 and 250 kilometres the score tapers, because very short sectors spend too much of the block time in taxi and climb. Beyond 700 kilometres it tapers again, and past 900 kilometres the model excludes the pair outright. Distances are great-circle, computed from airport coordinates.

The hard limits matter as much as the ideal band. They are what shrinks the search space to something honest: of every possible pairing of Indian airports with recorded traffic and known coordinates, exactly 2,743 pairs fall inside the envelope. The model ranks all of them, every run.

Question four: does it touch a base? (18%)

Aircraft, crews, and engineers sleep somewhere. For this carrier’s geography, that means the two bases. A pair touching either base scores 1.0. A pair touching neither scores 0.55, a penalty rather than an exclusion, because networks do grow tactical out-stations, but every orphan route drags positioning costs, duty-time complexity, and maintenance exposure behind it.

This is the only question that encodes anything specific about the operator, and even it uses only public knowledge: where the bases are. It is also, deliberately, the smallest weight. Demand should dominate geography, or the model degenerates into a map of where the aircraft already are.

What the model deliberately does not ask

The four questions are easier to defend once the rejected candidates are on record, because each omission was a choice rather than an oversight. The model does not score frequency potential, how many daily rotations a pair could sustain, because frequency is a schedule-design decision downstream of the market call, and folding it in would let an exciting schedule fantasy inflate a mediocre market. It does not score connecting traffic over the bases, because connect flows depend on bank timings the public record cannot see, and a regional network of this size lives or dies on point-to-point demand first. It does not score airport charges or fuel differentials between stations, real costs, but small enough relative to the demand question that modeling them from public tariff cards would add precision theatre rather than accuracy.

Hardest of all, it does not attempt a fare proxy. There was a tempting version of this model that estimated yields from distance and market-type heuristics, and it would have produced confident revenue numbers with nothing underneath them. Ranking on demand and saying so plainly beats ranking on invented revenue, because the first is an honest partial answer and the second is a complete answer that is quietly fiction. Every omission above follows the same rule: the model asks only questions the data can actually answer.

Why a linear model with visible weights

Four numbers, 34, 26, 22 and 18, invite an obvious challenge: why these? The honest answer is that they encode judgement, not regression. They were set by reasoning about the business, in the order the business reasons: size of opportunity first, momentum second, aircraft economics third, network gravity last. Anyone who disagrees can move the weights and re-run; the code makes that a thirty-second edit.

That re-runnability is the argument for simplicity. A gradient-boosted ranker fitted on the same file would produce a similar list and could never explain it to a planning committee, or survive one round of “what happens if we value growth more”. A linear scorer with four legible questions can be argued with, which is precisely what makes it usable in a room of professionals whose careers were built on this judgement. The model’s job is to hold the arithmetic steady while the humans argue about the weights. We wrote up the same philosophy for a different industry in our AI governance case study: systems that show their reasoning get adopted, and systems that ask for faith get shelved.

There is a deeper principle underneath, one we apply across every system we build. Derived judgements should be stored, inspected, and versioned, never regenerated opaquely on demand. A score of 91.2 for a route means something because the four sub-scores that compose it are written down beside it, and next month’s re-run can be compared against this month’s line by line. Intelligence that cannot show its working is not intelligence a regulated industry can act on.


The unglamorous work: making a public file tell the truth

Every data project has a stage nobody puts in the brochure. Ours consumed most of the build time, and skipping past it would misrepresent where the effort in this kind of work actually goes. The DGCA file is a gift, eleven years of monthly city-pair traffic, published free. It is also a file assembled by many hands over many years, and it behaves like one.

Start with names. The file spells Goa four different ways, depending on the year and the airport: plain Goa, Dabolim in a parenthetical, Mopa in two variants of its own, and, in the newest files, bare “Dabolim” with no mention of Goa at all. Treat those as separate cities and the state’s traffic quarters itself, wrecking every downstream number. That last variant deserves its own confession: an early build of this model carried aliases for only three of the four labels, and the bare Dabolim rows, over three million passengers a year of them, silently vanished from every Goa pair. The error surfaced only when a coverage check listed the busiest cities the model could not place and “DABOLIM” stared back from near the top of the list. Every number in this report post-dates that fix.

Goa is not special. Mumbai appears three ways once Navi Mumbai opens. Rajkot’s new airport arrives as “Hirasar (Rajkot)”. Hindon shows up as “Ghaziabad”. Tiruchirappalli gains a second “p” mid-decade. Vijayawada drops an “a”, Kolhapur carries a trailing space invisible to the eye and fatal to a string match, Agatti is sometimes an island and sometimes not, and Bengaluru shares a slash with Bangalore in one era of the file. The model carries an alias table that folds every variant to one canonical name, and that table was built the only way these tables ever are: by finding each mismatch after it silently bent a number that should not have bent, then adding a coverage check so the next one announces itself.

Then the parsing itself. Several city names contain commas inside quoted fields. Split the file naively on commas, the way a quick script does, and those rows shear apart into garbage that may or may not crash anything. The lesson is old but apparently needs re-learning once per project: a CSV is a format with rules, and only a real CSV parser knows them.

Then the artifacts, which are subtler because the arithmetic is technically correct. Extreme growth figures in this data come from two different illusions, and they demand different treatment. The first is the service-start artifact: Raipur–Visakhapatnam posts +602 percent on a prior-year base of 3,697 passengers, and Hindon–Varanasi posts +3,175 percent on a base of 3,011. Neither number means demand exploded. Each means a scheduled service started where none meaningfully existed, and the percentage is measuring the service, at its own induced demand, months into existence. The second illusion is nastier: before the alias fix described above, one Goa pair showed +7,123 percent growth that evaporated to +2 percent the moment the labels merged, because the “growth” was traffic migrating between spellings of the same airport. A model that cannot recognise these for what they are will confidently rank a route launch, or a renamed row, as a demand explosion.

Our handling is the base-500 floor described above, which mutes the worst of it, plus a flag rather than a fix for the rest: large growth on a recently thin base is labelled as a service-start artifact wherever it survives into output. We chose labelling over deletion deliberately. A new service growing into induced demand is genuine information about a market. It is just a different fact than organic growth, and the reader deserves to know which fact they are looking at.

One more boundary case earns a mention because it shows the model disagreeing with a route that exists. The carrier serves a coastal pair whose great-circle distance falls below the model’s 150-kilometre floor, so the model refuses to score it at all. The floor is economically right, sectors that short rarely cover their cycle costs on fixed-wing economics, and the route in question exists under a regional connectivity award, where viability gap funding changes the equation the model scores. We left the floor in place and noted the exclusion. A model that quietly bends its rules to flatter the observed network has resigned from its actual job.

The reason to dwell on all this goes beyond craft pride: the difference between a credible model and an embarrassing one lives almost entirely in this layer. The scoring logic of the four questions fits on an index card, and any competent analyst could rebuild it in an afternoon. What they could not do in an afternoon is know that Kolhapur carries a trailing space. Public-data work compounds: the aliases, floors, and flags built for this model now transfer to every Indian aviation question we touch next.


The test it had to pass first

Everything above is preamble to one moment: asking the finished model, which knows nothing about any carrier, which of the 2,743 pairs it likes best, and then checking where the carrier’s actual network landed in that ranking.

The protocol deserves one sentence of pedantry, because blind tests are only as honest as their sequencing. The four questions and their weights were written and locked first, from business reasoning alone. The list of the carrier’s routes was assembled separately, from public schedule information, and compared against the ranking only after the ranking existed. No weight was revisited after the comparison. A model tuned until it flatters the known answer proves only that its author can tune, which is why the freeze came before the look.

Here is the result, in full. The model scored and ranked all 2,743 pairs inside the ATR envelope. The carrier’s network placed as follows.

Route Rank of 2,743 Score Distance Passengers, trailing 12m
Agatti – Goa 6 95.8 538 km 38,619
Hubli – Hyderabad 9 94.0 414 km 55,493
Jalgaon – Pune 36 89.2 319 km 40,774
Goa – Jalgaon 38 89.0 649 km 41,267
Hyderabad – Rajahmundry 40 88.6 360 km 303,230
Hyderabad – Vijayawada 72 82.9 264 km 367,588
Goa – Pune 74 82.7 356 km 456,149
Bengaluru – Hubli 82 81.8 371 km 103,909
Bengaluru – Goa 317 51.3 483 km 1,526,595
Goa – Hyderabad 323 50.8 533 km 1,265,807
The first 100 ranks of 2,743 — the carrier's routes marked
69363840727482 rank 1 rank 100 of 2664
#317 Bengaluru–Goa · #323 Goa–Hyderabad sit beyond the first 100
Filled cells are routes the carrier flies. The two beyond the strip are the metro pairs that outgrew the turboprop band.

Eight of the carrier’s ten routes sit in the model’s top 100 of 2,743, two of them inside the top ten. Were the ranking random, the expected number of those ten routes falling in any top-100 is 0.36. The model was never told these routes existed as a network. It found them because the demand signature that made a veteran planning team choose them is written plainly into the public record, and the four questions are tuned to read exactly that signature.

Consider what each of those top ranks is actually saying. The island pair at rank 6 combines a captive market, zero surface-transport alternative, and a sector length in the heart of the turboprop band. Hubli–Hyderabad at rank 9 is a state-capital connection growing 27 percent on the exact natural size for a daily rotation. The pairs at 35, 37 and 39 are the classic regional play: state-capital and delta-city traffic, too thin for constant jet service, thick enough to fly daily. A planning team saw all of this with their own methods. The model saw it with four questions and a public file.

Now the two entries at the bottom of the table, because they look like failures and are the opposite. The two metro pairs rank 317 and 323, carrying 1.3 and 1.5 million passengers a year. If the model were a general-purpose route recommender, ranking million-passenger markets in the 300s would be an error. But the model scores turboprop opportunity, never importance, and pairs that size outgrew the 70-seat sweet spot years ago. The jets are already there. The correct behaviour for this scorer is to mark such pairs down, and it does, for precisely the stated reason: the size question caps out and reverses above the thin-route band. The carrier flies both pairs anyway, plausibly for network feed, brand presence, and aircraft positioning, which are real strategic reasons living outside a demand model’s jurisdiction. The model does not need to agree with every route to be useful. It needs to disagree legibly.

Three mild disagreements sit in the 72-to-82 band, and they repay a look at the sub-scores. Goa–Pune, at 456,000 passengers a year, is drifting past the top of the 70-seat band; the model docks its size score for the same reason it buries the metro pairs, just more gently. Hyderabad–Vijayawada and Bengaluru–Hubli lose points on growth, not size: solid markets whose trailing-year momentum was modest in this window. That is the model saying “sound route, maturing” or “sound route, not currently a wave”, which are defensible reads and, more importantly, inspectable ones. Anyone can open the sub-scores and see the sums they compare.

The distribution behind the table matters as much as the ranks in it. Plot all 2,743 scores in rank order and the curve falls steeply through the first hundred or so pairs, then flattens into a long, undifferentiated middle. The carrier’s routes cluster on the steep part, which is the meaningful placement: the model separates its top few dozen pairs decisively, and the network sits inside that separated group rather than scattered through the plateau where rank differences mean little. A rank of 9 on this curve is a strong claim. A rank of 900 versus 1,000 is noise, and the model’s users are told so.

A word on what this validation does and does not establish, because it is easy to over-claim and we would rather under-claim. Formally, the test shows that the model’s ranking function, built and weighted before any comparison against the carrier’s network, assigns top-4-percent ranks to eight pairs the carrier independently chose. The weights were set by business reasoning, and were locked before validation. What the test does not establish is predictive power over profitability. A route can be demand-perfect and still lose money on fares, and the model says nothing about fares, a limit this study takes seriously in its own section.

What the test buys, in practical terms, is standing. When this model subsequently claims that an unserved pair resembles the carrier’s best routes, that claim inherits credibility from the blind rediscovery. The claim has a specific, checkable meaning: the pair scores the way Agatti–Goa and Hubli–Hyderabad score, on the same four questions, in the same public data. That is a different kind of statement than a consultant’s “we see potential here”, and the difference is the entire point of building the instrument before making the argument.


What it found: three pairs the carrier does not fly

With the validation on the table, the same ranking can be read the other way. Strip out every pair the carrier already serves, and look at what remains at the top. This study shows three of the resulting pairs, chosen because each tells a different story about how the model argues. Several more cleared the same scoring bar, and we hold them back deliberately: a ranked list handed over wholesale invites armchair rebuttal, while a shortlist worked through in a room, beside an operator’s own numbers, becomes a plan.

Route Rank of 2,743 Score Distance Passengers, trailing 12m Growth
Calicut – Hyderabad 8 94.7 728 km 141,183 +42%
Hyderabad – Kolhapur 28 91.2 445 km 51,615 +16%
Bengaluru – Kannur 56 85.2 274 km 113,848 +24%

Calicut – Hyderabad: the wave

The largest and fastest-growing of the three, and a top-ten pair nationally. Just over 141,000 passengers moved between these cities in the twelve months to July 2026, nearly twelve thousand a month, and the market grew 42 percent year on year from a healthy base, which rules out both artifact families. The demand logic is legible from the ground: Kerala’s third metro, dense with students, IT workers and Gulf-returning families, connecting to the Deccan’s technology capital. The score anatomy: full marks on size, 34 of 34. Full marks on base connection, 18 of 18, since one end is the carrier’s Deccan base. Two deductions: growth takes 23.9 of 26, because 42 percent sits below the 50-percent cap, and range takes 18.8 of 22, because 729 kilometres sits past the 700-kilometre ideal band. Total, 94.7, rank 8 of all 2,743.

Calicut – Hyderabad · rank 8 of 2,743
Right size34 of 34
Growing23.9 of 26
Range18.8 of 22
Base18 of 18
Score94.7
141,183 passengers in the trailing twelve months, up 42% year on year. Deductions: the 728 km sector past the comfortable band, and growth below the 50% cap.

The strategic read: a workforce-and-family corridor between Kerala’s third metro and the Deccan’s technology capital, growing at a rate that suggests the market is being discovered rather than merely served. The 728 kilometres push an ATR toward the edge of comfort, roughly two hours in the air, and the score already carries that penalty honestly rather than hiding it.

Hyderabad – Kolhapur: the textbook

Rank 28, and the purest expression of what the model is for. About 4,300 passengers a month, which is almost exactly the natural size for one daily 70-seat rotation. Growth of 16 percent on a clean multi-year base: organic, unspectacular, real. Sector length of 445 kilometres, dead centre of the ideal band, 22 of 22. One end at the carrier’s Deccan base, 18 of 18. The only deduction is growth, 17.2 of 26, because 16 percent is solid rather than explosive. Total, 91.2.

Hyderabad – Kolhapur · rank 28 of 2,743
Right size34 of 34
Growing17.2 of 26
Range22 of 22
Base18 of 18
Score91.2
About 4,300 passengers a month, the natural size for one daily 70-seat rotation. The only deduction is growth: 16% is solid rather than explosive.

No individual number here is exciting, which is exactly the point. This is the kind of pair a human screen misses: too small to make a conference presentation, growing too gently to make a headline, connecting a base to a city with no direct service from it. The model surfaces it because arithmetic does not get bored. Pairs like this are where a scoring instrument pays its rent, finding the quiet route that fits the aircraft like a glove while attention is on the loud ones.

Bengaluru – Kannur: the demand that scores despite the map

Rank 56, and the most instructive of the three, because it ranks there while carrying the model’s biggest penalty. Neither end touches a base, so the pair takes the full base deduction, 9.9 of 18. It ranks 56th of 2,743 anyway, on the strength of everything else: nearly 114,000 annual passengers, perfectly sized at 34 of 34; growth of 24 percent on a six-figure base, 19.2 of 26; and a 274-kilometre sector inside the ideal band, 22 of 22. Total, 85.2.

Bengaluru – Kannur · rank 56 of 2,743
Right size34 of 34
Growing19.2 of 26
Range22 of 22
Base9.9 of 18
Score85.2
Neither end touches a base, so the pair carries the model's full base penalty and ranks 56th of 2,743 anyway. The demand is that strong.

When a pair absorbs the largest structural penalty in the model and still outranks 2,609 of 2,743 candidates, the demand is not marginal. The model is saying: this market is strong enough to justify thinking past the current base map, whether that means tagging the pair onto existing flying or treating it as the seed of a future out-station. That decision belongs to humans with cost sheets. The model’s contribution is to make sure the question gets asked.

Reading a score anatomy

The three anatomies above follow a pattern worth making explicit, because it is the skill the instrument teaches its users. A high total score says little on its own; the information lives in where the points were lost. Calicut–Hyderabad loses only on range, so the operational question it poses is about sector economics at the edge of the envelope: fuel burn, crew duty fit, and whether the schedule can hold two-hour rotations. Hyderabad–Kolhapur loses only on growth, so its question is about ceiling: does a 16-percent market keep absorbing seats, or does one daily rotation saturate it? Bengaluru–Kannur loses only on base fit, so its question is purely logistical: what does serving this pair cost in positioning and crew arrangements, given nothing sleeps at either end?

Three routes, three completely different due-diligence agendas, all read directly off which sub-score paid the penalty. A single opaque score, however accurate, could never do this. The decomposition converts a ranking into a work plan, and that conversion, more than the ranking itself, is what a planning team actually buys when it adopts a legible instrument.

The shape of the finding

Put the three pairs on a map and a pattern appears that no single score states: all three sit inside the geographic footprint the fleet already serves. None requires the carrier to go somewhere new. They ask the network to catch demand already flowing through its own neighbourhood, past aircraft that already sleep nearby. Growth plans usually imagine expansion as reaching outward; the model’s first suggestions are all about density within reach.

The network, and the three pairs the model likes
Agatti Bengaluru Calicut Goa Hubli Hyderabad Jalgaon Kannur Kolhapur Pune Rajahmundry Vijayawada
Flown todayScored, not flown
Solid lines are flown today. Dashed lines are scored, not flown. Positions are approximate great-circle geometry, not a projection.

That is also the finding’s built-in caveat. The model sees demand, and demand within reach is necessary but far from sufficient. Whether any of the three pairs survives contact with fares, competitor frequencies, and slot times is a question the public record cannot answer, which brings this study to its most important section.


Presenting the analysis: a report the reader flies through

An analysis that is right but unread changes nothing, so the presentation layer got the same design attention as the model. The findings shipped as a single self-contained web page, and its construction is part of this case study because the choices generalise to any analytical deliverable meant for a senior operator rather than an analyst.

The page is built around one continuous journey. A map of western and southern India sits fixed behind the text, and as the reader scrolls, a camera flies over it: the network draws itself route by route in the order the model ranked it, and each of the three unserved pairs is flown as its own leg, an aircraft tracing the dotted line from city to city while the numbers for that pair arrive alongside. Scrolling is the only interaction the reader must learn. Everything else, the score anatomies, the rank strips, the theme, is optional depth.

Four presentation decisions did the persuasive work.

Validation before opportunity, always. The page’s structure forces the order of the argument: the reader meets the model’s method, then the proof that it rediscovered the existing network, and only then any suggestion. A reader who has just watched a blind model rank their industry’s known-good decisions in the top four percent extends provisional trust to the next screen. The same three pairs, presented first, would read as one more vendor’s opinion.

Every score opens into its arithmetic. Each route card unfolds, on a tap, into the four sub-scores with the exact points earned and forfeited: 18.3 of 22 on range, 9.9 of 18 on base fit. Nothing is presented as a verdict that cannot be decomposed. This is the interface expression of the model’s whole philosophy, and readers notice the difference between a number they can open and a number they must accept.

The misses are annotated, not hidden. The validation table on the page includes the two metro pairs at the bottom of the ranking, with a note explaining why a turboprop-opportunity scorer should rank million-passenger markets low. Leaving the worst-looking rows in, with their reasoning, bought more credibility than the top-ranked rows did on their own.

The limits get their own chapter, on the page, unprompted. Fares, competitor schedules, slots, and growth artifacts are stated in the report itself, in plain language, before any reader can raise them. In every conversation since, the limits section has been the part sophisticated readers mention first, and approvingly.

The choice of a live page over a PDF was itself an argument. A slide deck freezes an analysis at the moment of export and asks to be believed; a page whose tables respond, whose anatomies open, and whose map flies invites the reader to poke at the thing, and poking is how sceptics convert. The medium also matches the model’s own promise: something built to re-run monthly should arrive in a form that visibly could, rather than as a snapshot with a date already going stale in the footer.

Two smaller courtesies rounded it out. The page is one file with no external dependencies, so it loads anywhere, works offline once opened, respects dark mode and reduced-motion preferences, and prints cleanly to paper for the reader who wants to scribble on it. And it was published quietly, unindexed, addressed to its intended reader rather than to search engines, because a demonstration built for one operator should behave like correspondence, courteous even in its metadata.


What the model cannot see, stated before anyone asks

A model earns trust by the limits it declares, so this section is the longest honest list we can write. Four things are invisible to a scorer built on the DGCA city-pair record, and each one is capable of reversing a route decision on its own.

Fares

The DGCA publishes how many people flew. It does not publish what they paid. That gap is not a rounding error; in thin-route economics it is frequently the whole question. Two pairs can carry identical passenger volumes while one sustains fares that cover a turboprop’s trip cost and the other survives on distressed pricing that would drown any new entrant. Yield varies by season, by day of week, by booking window, and by the mix of leisure, business, and obligation travel, and none of that variance leaves a trace in a passenger count.

The consequence for reading this study is direct: every rank in the findings above is a statement about demand, never about revenue. A pair the model loves could be a fare graveyard. The reverse error also exists, and is subtler: a pair the model scores as merely good might clear exceptional yields, because captive markets with no rail alternative pay for their geography. The island route at rank 6 exists partly on that logic. Distinguishing the two cases requires fare data, and fare data lives inside airline revenue-management systems and the distribution channels they feed, behind exactly the wall a public-data model respects.

Competitor capacity

The model sees that 141,000 passengers flew between Calicut and Hyderabad. It cannot see who carried them, at what frequency, on what aircraft, or with how much unsold capacity. A pair whose demand is growing at 42 percent while incumbent capacity grows at 20 is an invitation. The same demand curve with incumbents adding seats faster than passengers is a knife fight dressed as an opportunity. The carrier-level file the DGCA publishes gives airline market shares in aggregate, but route-level schedules and seat counts belong to schedule databases and the operators themselves.

This limit interacts with the growth question in a way any professional reader will spot at once. High growth attracts capacity, and capacity, once landed, converts a wave into a share war within two seasons. A demand model flags the wave. Only a competitive overlay can say whether the wave still has room on it. Building that overlay from published schedules is possible, and it is the first improvement we would make with a week of an operator’s time; it simply cannot be built from the passenger record alone, and pretending otherwise would poison the numbers that are sound.

We tested the obvious shortcut, and the test is worth reporting because it failed in an instructive way. An openly published scraped archive of Indian domestic schedules covers fourteen airlines and 99 cities across seven years, and it does not contain the carrier this study models at all. An entire operating airline, absent from the dataset everyone would reach for first. That is the case against scraped competitive data in one sentence: it fails silently, and it fails on exactly the small operators a regional analysis cares about.

Read purely as an illustration, though, the same archive sketches what a real overlay would say. On all three pairs highlighted in this study, its 2025-valid entries show a single mainline operator holding the schedule at roughly daily to four-times-daily frequency, while the demand beneath grows at 16 to 42 percent. One incumbent, jet-sized equipment, demand outpacing anything a single schedule can absorb: that is the entrant’s pattern, and it is precisely what the overlay exists to detect. We publish the sketch with its provenance stated and its numbers unaudited, because the point is the method’s shape, and the honest version of it runs on operator-grade schedule data or nothing.

Slots and ground reality

A pair can be perfectly sized, growing, in range, and touching a base, and still be unflyable at any commercially sane hour, because the airport at one end has nothing to offer but a 14:40 arrival into a business market or no parking stand at all. Slot scarcity at Indian metros is a live constraint, and it binds hardest exactly where regional carriers want to connect: the big-city end of a thin pair, where feed and self-connection value is highest. Beyond slots sit the quieter ground constraints, counter space, ground-handling contracts, night-parking, fuel availability at smaller fields, none of which appear in any public dataset at usable fidelity.

The honest framing is that the model ranks markets, and markets are only two-thirds of a route. The remaining third is operational permission, and operational permission is negotiated, not computed.

Growth artifacts, restated as a warning label

The data-work section described the service-start artifact from the builder’s side; it belongs in this list from the reader’s side, because it is the one limit that actively manufactures false confidence. A pair posting three-digit growth may be a discovered market, or it may be one airline’s year-old schedule generating its own induced demand, and the difference decides whether a second operator starves. The model mutes bases under 500 passengers and flags recent-thin-base growth, but muting and flagging are mitigations. The cure is time: watching whether the pair’s third and fourth semesters hold the level its first two claimed. Arithmetic can wait patiently. So should capital.

Why publishing the limits is the strategy

There is a school of pitch-craft that says lead with strength and let the gaps surface in diligence. We hold the opposite view, and this study is built on it. Every limit above defines, precisely, a dataset the model’s next version needs, and every one of those datasets lives inside an operator: fare performance, competitive intelligence, slot portfolios, route-level maturation curves. Declaring the limits is therefore identical to publishing the collaboration agenda. A reader inside an airline can look at this section and know, to the column, what would happen if their numbers met this instrument.

That is the deliberate design of the whole exercise. The public-data model is complete as a demonstration and incomplete as a decision tool, and the incompleteness points at exactly one door. Handing an operator a finished answer would be both wrong, for every reason in this section, and strategically empty. Handing them a validated instrument with four labelled input sockets is an invitation with its own evidence attached.


Running it as an instrument, not a report

A market study is an event. An instrument is a habit. The difference decides whether analytical work compounds or evaporates, and it is worth spelling out what the habit looks like with this model, because the operational profile is the least glamorous and most persuasive part.

The model is a few hundred lines of plain Python. It ingests two CSV files fetched from the public mirror, runs in seconds on a laptop, and prints its full ranking with every sub-score exposed. There is no service to subscribe to, no dashboard to license, no vendor in the loop. When the DGCA publishes a new month, the fetch-and-run cycle takes under a minute, and the June ranking can be laid beside the May ranking line by line. A planning meeting that happens monthly can open with a diff: which pairs moved, which sub-scores moved them, and whether the movement is signal or artifact.

Picture the ritual concretely. The July file lands. The run takes a minute. The diff shows Calicut–Hyderabad’s growth sub-score easing from 26.0 toward 24 as the explosive year laps itself, Hyderabad–Kolhapur holding steady, and a pair nobody was discussing creeping from rank 90 to rank 60 on two consecutive months of quiet gains. Ten minutes of a planning meeting, once a month, and the network’s peripheral vision stays swept. No standing analyst assignment, no quarterly study commissioned, no deck.

That cadence changes the nature of the questions a team asks. A static study invites a verdict: accept or rebut. A monthly instrument invites hypotheses: if this corridor’s growth holds above 20 percent for two more quarters, it crosses our threshold; watch it. Route selection becomes less like litigation over a consultant’s conclusions and more like a watchlist with tripwires, which is how experienced planners already think. The instrument does not replace the thinking. It gives the thinking a memory, and a memory that survives personnel changes.

Versioning the judgement matters as much as versioning the code. The four weights are an explicit, dated statement of belief: this is how much we value size against momentum against range against geography, as of this quarter. When the belief changes, because a new base opens or the fleet adds a longer-legged type, the weights change in one visible place, and every historical ranking can be re-run under the new belief for comparison. Institutional judgement usually lives in heads and hallway consensus. Here it lives in four numbers under version control, arguable by anyone, auditable by anyone, and that migration of judgement from folklore into inspectable artifacts is, in our experience across industries, worth more than the model’s own output. It is the same principle we apply when designing AI operating structures for enterprises: the thinking stays human, the arithmetic becomes infrastructure, and nothing derived is ever regenerated opaquely when it can be stored, dated, and signed instead.

The economics deserve one plain paragraph. This entire build, data acquisition, cleaning, modeling, validation, and the interactive report, was executed by a small team in days, on free public data, with no procurement cycle. That is not a boast about the team; it is a fact about the era. The marginal cost of turning a public record into a validated decision instrument has collapsed, and any analytical claim that used to justify a six-figure study now has to compete with what a competent group can build and validate in a week. Organisations that internalise this reprice their questions accordingly. The scarce input is no longer analytical horsepower. It is knowing which questions are worth instrumenting, and that knowledge still walks around in the heads of operators.

What would change with an operator inside the loop is worth stating concretely, because it converts the limits list from a confession into a roadmap. Fare data would turn the demand ranking into a contribution ranking, the difference between “people fly this pair” and “this pair pays for the aircraft”. Schedule data would add a saturation overlay, separating open waves from crowded ones. Slot and ground constraints would prune the list to the flyable subset before anyone falls in love with a number. And the operator’s own route-maturation history, how its young routes actually ramped, would calibrate the growth question against ground truth no public file contains. Each addition is a bounded, weeks-not-quarters piece of work, because the chassis they bolt onto already runs.


The questions a network planner would ask

This study will be read, if it is read at all, by people who have spent careers making the decision it models. They will have objections, and the objections are good ones. Here are the six we would expect from the sharpest reader in the room, answered the way we would answer them across a table.

”Our team already knows all of this.”

For the ten routes in the validation table: yes, demonstrably, and that is the finding. The model’s agreement with the existing network is presented as evidence for the team’s judgement, never as a substitute for it. The claim being made is narrower and, we think, harder to dismiss: no team of any size re-reads 2,743 pairs every month with equal attention. Human screening is priority-driven, and priority is set by what is already visible, which is how a Hyderabad–Kolhapur sits unexamined at rank 28 while attention goes to louder markets. The instrument does not know things the team does not know. It attends to things the team has no time to attend to, uniformly, monthly, without fatigue. Knowing and attending are different capacities, and networks are lost in the gap between them.

”Why should we trust weights you made up?”

You should not trust them; you should argue with them, and the model is built so that the argument is cheap. The four weights are one defensible codification of standard network-planning logic, and their defence is in the method section above. But the deeper answer is that the ranking is remarkably stable under reasonable weight changes. Swap size and growth, at 26 and 34, and the validation routes stay in the top decile, because pairs that score well on this data score well on several questions at once. Weight sensitivity is a fifteen-second experiment here, and running it live in front of a sceptic, watching the good pairs stay good while the borderline pairs shuffle, has convinced more people than any argument in prose. A model you can stress-test in the meeting is a different object from a model you are asked to believe.

”Why not machine learning?”

Because the constraint that binds this problem is trust, and trust here flows from legibility. A learned ranker fitted on eleven years of this file would likely score pairs a few percent “better” by some offline metric, and would be unusable in the only setting that matters: a room where a decision-maker asks “why this pair and not that one” and needs an answer that survives cross-examination. Four questions with visible arithmetic answer that question by construction. There is also a statistical honesty issue: with under 700 in-envelope pairs and a target as noisy as route success, a flexible learner has more capacity than the data has lessons, and what it would learn most fluently is the incumbents’ historical biases. The linear scorer is not a compromise awaiting an upgrade. At this data scale, it is the correct tool.

”What about seasonality?”

Handled structurally rather than cleverly. Every level the model uses is a trailing twelve-month sum, and every growth figure compares one complete year against the previous complete year. Each window contains exactly one monsoon, one winter peak, one summer trough, so seasonal shape cancels out of both the size and growth questions. What the model deliberately does not do is score within-year shape, a pair that fills aircraft in December and empties them in July may deserve seasonal service rather than year-round, and that distinction needs monthly-resolution analysis the current scorer skips. The monthly curves exist in the same file, so the seasonal-shape overlay is another bounded extension, not a redesign.

”The data is a year old by the time it matters, and it counts flown passengers, not demand.”

Both true, and worth being precise about. The DGCA record runs about two to three months behind the calendar, and a trailing-year window smooths away the newest signal by design; this instrument reads tides, not weather, and a schedule decision made for the next season is a tide decision. The second point is the deeper one: the file records passengers carried, which is demand as filtered through the capacity airlines chose to offer. A pair with no service shows zero passengers however many people wanted to fly it, so true white-space, city pairs with demand and no flights at all, is structurally invisible here. That blind spot suits this model’s user, an operator scaling within a known region, better than it suits a greenfield planner: the pairs it surfaces are markets already proven by someone’s aircraft, at volumes the incumbent offering has not exhausted. Estimating unserved-pair demand needs gravity modeling on population and economic data, a different instrument with its own uses and its own, larger, error bars.

”Eight of a hundred could be luck. How do we know the test is strong?”

By counting. The carrier’s network occupies ten slots among 2,743 ranked pairs. If the model’s ranking were random, the expected number of those ten falling in any top-100 is 0.36. Eight landed there, two of them in the top ten. The probability of eight or more of ten specific pairs falling in a random top-100 of 2,743 is astronomically small; this is not a marginal signal that needs a p-value dressed up to look respectable. The honest challenge is different: the model’s questions encode turboprop logic, and the carrier is a turboprop operator, so of course a turboprop-shaped sieve catches turboprop-shaped routes. To which we say: yes, exactly. The claim was never that the model contains magic. The claim is that the operative logic of a specific, difficult business decision can be written down in four questions, run against a public file, and shown to reproduce expert choices. The demystification is the product.

”What would it take to trust this for an actual launch decision?”

More than this model, and the model says so itself; the four stated limits are the checklist. Concretely, we would want the demand rank confirmed by a fare read on the pair, a competitive overlay showing incumbent capacity and its trajectory, a slot answer at both ends at commercially useful hours, and two further quarters of the growth signal holding. The model’s role in that sequence is triage, spending expensive analytical attention only where cheap public arithmetic already found a pulse. No instrument of this kind should ever be the last word before capital moves. Its job is to be the first word, reliably, every month, so that the expensive investigation starts from a ranked shortlist instead of a blank map.


The instrument, one week later: the whole record, and the questions it can now answer

The study above ran on one file: the city-pair record, aggregated by an open mirror. In the week after it was written we went to the source and took everything the DGCA publishes under Aviation Data and Statistics: the same city-pair tables back to April 2015 (two months the mirror lacks, recovered from the DGCA’s own PDFs), every airline’s monthly operating statistics from 2009 (departures, block hours, seat-kilometres, load factor, cargo), the international quarterly tables from 2015, and 144 monthly traffic reports with on-time performance, cancellations and complaints. All of it now lives in one warehouse with canonical city names and seven reconciliation checks. The check that matters most: the city-pair totals and the airline totals, two separate DGCA publications counting the same passengers, agree to the passenger in 135 of 135 months.

The warehouse edition, trailing twelve months to July 2026
136
Months of city-pair data, no holes
April 2015 to July 2026; two months rescued from PDF
135 / 135
Months where the two DGCA tables reconcile
City-pair totals equal airline totals to the passenger
8 of 10
Carrier routes in the blind top 100, unchanged
Now of 2,743 scored pairs; 79 more cities carry coordinates
144
Monthly traffic reports parsed
On-time performance, cancellations, complaints, 2014 to 2026

With the airline tables in the same place as the city pairs, the model stops being one screen and becomes a set of instruments. Each is stated with its window and its own measure of error, because that is the only way a network planner will let a number into the room.

The operators, measured in their own unit

The monthly operating statistics let the industry be read in the unit a regional CEO actually manages: block hours, departures, seat-kilometres per hour. Over the twelve months to July 2026 the industry flew at an 85.1 percent passenger load factor. The regional operators combined flew at 71.9. Inside that group the spread is the story: FLY91 at 76.4 percent (up 2.1 points on the prior twelve, on departures that grew 89 percent), Star Air at 74.3, IndiaOne Air at 65.1, Alliance Air at 64.2 and down 10.3 points. Regional aircraft fly 300 to 540 kilometre stages at 1.2 to 1.4 block hours per departure; the narrowbody operators fly 880 to 1,180 kilometre stages. Same country, two different businesses, now visible in one table.

Passenger load factor, scheduled domestic, twelve months to July 2026
Akasa Air
92.3%
SpiceJet
85.6%
IndiGo
85.2%
All carriers
85.1%
Air India Express
84.6%
Air India
82.9%
FLY91
76.4%
Star Air
74.3%
IndiaOne Air
65.1%
Alliance Air
64.2%
Source: DGCA monthly traffic and operating statistics (ICAO Form A). LF = passenger-km / available seat-km.

How a new route actually ramps

The service-start artifact described earlier had a positive twin waiting in the same file. If the record shows every route that was ever born, it shows how routes grow up. We identified every true new city pair since April 2016, a pair with no earlier traffic that reached 500 passengers in one of its first three months, with the pandemic restarts excluded: 395 launches, 145 of them in the turboprop class of 2,000 to 8,000 passengers a month. Their median trajectory is the benchmark no planning team had, because no single team launches enough routes to compute one.

The median turboprop-class route carries 3,863 passengers a month at month three, 3,046 at month twelve, and 1,594 at month twenty-four. Sixty-five percent are still flying at month twelve and 70 percent at month twenty-four (some routes pause and return). The typical new route peaks early and erodes; a route that is still climbing in its second year is doing something unusual. Laid against that curve, the carrier’s own 2024 launches sit above the median through month twenty-four, which is the kind of finding a public record can hand a planning team as evidence rather than as praise.

Median monthly passengers, turboprop-class launches since 2016 (145 routes)
Month 11921 of 3863
Month 33863 of 3863
Month 63762 of 3863
Month 123046 of 3863
Month 241594 of 3863
Bars show the median route's volume at each month as a share of its month-three peak. Share of routes still flying: 100 percent at month one, 65 at month twelve, 70 at month twenty-four.

Trained on the 332 launches with two observable years, a classifier that sees only what is known by month three (early volume, distance, catchment population, endpoint throughput, launch month, whether a metro is touched) predicts survival to month twenty-four with a cross-validated AUC of 0.73 against a base rate of 61 percent. Modest, honest, and enough to rank this year’s launches by the odds the record gives them.

Latent demand: what a pair should carry

The screen ranks pairs that already show demand. A gravity model asks the prior question: given two cities’ catchment populations, their airport throughput and the distance between them, how many passengers should fly, and where is the gap? Fitted on 613 served pairs, the model explains two-thirds of the variance out of sample (cross-validated R-squared 0.67) with a distance elasticity of minus 0.67 and a penalty on pairs under 300 kilometres, where road and rail carry the trip. The residual is the finding. Hyderabad–Nagpur carries 109,000 passengers a year against a potential near 378,000; Bengaluru–Tirupati 55,000 against 275,000. Among pairs with essentially no service, Bagdogra–Patna, Mangalore–Pune, Ranchi–Varanasi and Goa–Mangalore each show a potential above 100,000 passengers a year on the model’s terms. Potential is not a booking; it is where the arithmetic says to look first.

A forecast that states its error

Every pair with more than a thousand passengers in the last year now carries a twelve-month forecast, from gradient-boosting models fitted per horizon on the pair’s own history and calendar. The number that matters is not the forecast but the backtest: trained to July 2025 and asked to predict the following twelve months, the model’s weighted absolute error across 666 pairs was 18.3 percent against 31.5 percent for a seasonal-naive rule; on trunk routes 12.4 percent, on turboprop-class pairs 27.7 against 63.1; on the national total 4.7 percent against 14.3. Looking forward from July 2026 it puts the next twelve months at 163.2 million passengers against 167.5 million in the last twelve, a soft year, with bands attached to every pair. A new competitor, a fare shock or a grounding are not in the features, and the model says so.

What exits leave behind

The airline tables reach back to 2009, far enough to watch three carriers leave. Jet Airways held 10.8 percent of scheduled domestic passengers in the six months before it stopped in April 2019; in the six months after, IndiGo gained 4.3 points of share, SpiceJet 2.6, Go First 2.3, each on capacity growth of 15 to 45 percent year on year. Go First’s own 7.4 percent in 2023 went 7.1 points to IndiGo. The industry level is as far as public data reaches, since the DGCA does not publish carrier-by-pair traffic; but the shape of redistribution is now a measured fact rather than a recollection.

Service quality, from the regulator’s own reports

The 144 monthly traffic reports add the second dimension a passenger feels. Across calendar 2025, on-time performance at the six metro airports averaged 81.6 percent for IndiGo, 78.4 for Akasa Air, 76.4 for the Air India group, 58.3 for Alliance Air and 58.1 for SpiceJet. The overall cancellation rate in June 2026 was 0.63 percent. Read beside load factors and share, these numbers turn an operator table into an operator portrait.

Ask the data

All of this sits behind the panel at the end of this page. It is the warehouse as a conversation: an analyst that answers only from queries it runs in front of you (every answer opens to show them), states the window and the source with each figure, and refuses to judge viability or profitability because the data cannot. Ask it what a pair carried, how a launch compares with the ramp benchmark, which regional operator flies fullest, or what it cannot see. The same tools are exposed as an MCP server, so a planning team can put the record and the models inside whatever assistant they already use.


The pattern, beyond aviation

Strip the propellers off this study and a generic method remains. It has five steps, and each transfers to any industry where a recurring, high-stakes selection decision meets a public record.

Find the public spine. Every regulated industry publishes more than its participants habitually use: drug approvals and shortage lists, port throughput, tender awards, corporate filings, sanction registries, land records, trade statistics. The DGCA city-pair file is unusually clean-hearted, eleven years, monthly, universal, but it is not unusual in existing. The first hour of this method is always the same question: what does the regulator already count?

Encode the judgement, small. The four questions came from reasoning about how a network planner thinks, written down and weighted. Every operating domain has an equivalent four-to-six questions that its veterans apply instinctively: what makes a good tender to bid, a good catchment for a clinic, a good SKU to range, a good account to call first. The discipline is keeping the model small enough that every score can be defended aloud, because the moment a score cannot be explained, the room reverts to instinct and the instrument dies.

Validate blind against known decisions. This is the step that separates the method from a spreadsheet with opinions, and it is available in any domain with an observable track record. Before the model may recommend, it must rediscover: rank the choices the organisation already made, without being told what they were. Pass, and every subsequent suggestion inherits earned credibility. Fail, and the model was wrong in private, at a cost of nothing, which is the cheapest place a model can fail.

Declare the limits as a data agenda. Whatever the public spine cannot see, fares here, margins or capacity or intent elsewhere, becomes a named list, and the list doubles as the integration plan for the owner’s internal data. Limits stated up front convert scepticism into scoping.

Run it on a cadence, and keep the judgement versioned. The instrument re-runs when the record updates, its weights live under version control as dated statements of belief, and its outputs are stored beside the reasoning, so that every future disagreement is a diff rather than a debate about recollections.

To make the transfer concrete, run the method against a different industry in one paragraph. A pharmaceutical distributor deciding which hospital accounts to pursue faces the same shape of problem as the route planner: thousands of candidates, a few dozen viable, judgement concentrated in senior heads. The public spine exists, in drug procurement tenders, hospital bed registries, and regulatory licensing data. The house judgement compresses to a handful of questions: is the account the right size for our logistics footprint, is its purchasing growing, is it within our cold-chain radius, does it sit near accounts we already serve? The blind validation is available, because the distributor’s current book of accounts is a set of historical decisions the model must rediscover before it may recommend. Same instrument, different nouns. Nothing in the aviation case depended on aviation; it depended on a public record, a compressible judgement, and a checkable track record, and those three ingredients are common.

None of these steps requires heroic technology, and that observation is the quiet thesis of this study. The current excitement around AI in the enterprise concentrates on generation, chatbots, copilots, document drafting, and most organisations have now bought or built the easy layer. The compounding value sits a level up, in decision instrumentation: taking the selections that senior people repeat under uncertainty and giving each one a public-data spine, an explicit scoring of the house judgement, a blind validation, and a monthly cadence. Aviation route selection happens to demonstrate the pattern vividly, because the public record is generous and the decisions are famous. The pattern itself is portable to any desk where someone repeatedly asks “which one next?”


Reproducing this study

Everything in this study can be verified from a laptop, and a claim of reproducibility should come with the actual recipe, so here it is.

The data is the aggregated domestic city-pair file from the india-aviation-traffic open-source mirror of DGCA monthly statistics, two CSVs totalling about 4.5 MB. The model is a single Python script with no dependencies beyond the standard library. The run takes seconds and prints two blocks: the validation table, the carrier’s known network as ranked blind, and the top unserved opportunities with their sub-scores.

The parameters that define the model, in one place: envelope 150 to 900 kilometres with an ideal band of 250 to 700; size sweet spot 2,000 to 25,000 passengers per month with proportional decay outside it; growth measured as trailing twelve months against the prior twelve with a 500-passenger base floor and neutral scoring below it; base cities weighted at full score for a touch and 0.55 otherwise; weights 34, 26, 22, 18. Distances are great-circle from published airport coordinates. City-name aliasing is applied before any aggregation, and the alias table is part of the code, not a preprocessing step someone has to remember.

Anyone re-running against a later DGCA month will get different numbers, which is the point of an instrument. The validation claim is dated: as of the twelve months ending July 2026, the ranks are the ones printed in the validation table, and the model that produced them was frozen before the comparison. If a later month moves a rank, that is the market speaking, and the correct response is curiosity about the sub-score that moved.

The warehouse edition adds a second recipe. The full DGCA export, the parsers, the reconciliation gates, the models and the chat tools live in one repository; aviation build, aviation gates, aviation models and aviation report rebuild everything from the raw files in about twenty minutes, and aviation mcp exposes the same four tools (schema, read-only SQL, model outputs, document search) to any MCP-capable assistant.

Adapting the study to a different operator is a short checklist rather than a rebuild. Replace the base set with the operator’s own bases, which changes one line. Replace the envelope parameters if the fleet differs, a 90-seat regional jet stretches the range band and raises the size band, and both edits are two numbers each. Re-examine the alias table against any city the new network touches that this one did not, because the naming quirks are the one part of the work guaranteed to resurface. Then run the same blind validation against the new operator’s known network before reading a single opportunity, and hold the same rule we held: a model that cannot rediscover the existing network has no business suggesting additions to it. The rule costs nothing to enforce and is the entire difference between an instrument and a slide.


A last word on what this study asks of its reader. If you run a network, the invitation is explicit in the limits section: the model’s four blind spots name the four datasets that would complete it, and the completion is weeks of work, on your premises, against your numbers. If you run something that has nothing to do with aircraft, the invitation is the pattern section: somewhere in your operation is a recurring selection decision, a public record that bears on it, and a handful of questions your best people already ask instinctively. The distance between those raw materials and a validated instrument is far shorter than the analytics industry has trained you to expect.

What did this exercise prove, in the end? Three things, in ascending order of consequence.

It proved that a decade of Indian regional aviation decisions is legible in public data, which is a compliment to the DGCA’s publishing and to the planners whose choices left such a clean signature. It proved that the judgement behind a high-stakes recurring decision can be written down small, validated blind, and turned into an instrument that costs nothing to re-run, which is a statement about method available to any industry with a public record. And it proved, to us most of all, that the fastest way to earn a serious operator’s attention is to do real work on their problem before asking for their time, publish the arithmetic, and state the limits before being asked.

The model is running. The record updates monthly. And 2,743 city pairs are now being watched by something that never gets tired of watching, waiting for the numbers only an operator holds to finish the picture.

Frequently asked questions

Can a route model built only on public data be trusted?+

It has to earn trust through blind validation. Before suggesting anything, this model was asked to rank all 2,743 feasible city pairs without being told which routes the carrier flies. It placed eight of the carrier's ten routes in its top 100, against a random expectation of well under one, which means the ranking logic reproduces expert network-planning decisions from public data alone.

What data does the route model use?+

One public source: the DGCA's monthly record of domestic passengers carried between every pair of Indian cities, from April 2015 to July 2026. No fares, no schedules, and no carrier-internal data of any kind are inputs.

What can the model not see?+

Fares, competitor capacity, and airport slots. It ranks demand, never profitability. Each of those blind spots names a dataset that only an operator holds, which is exactly where the model's next version would come from.

Ask the data

The warehouse behind this report, as a conversation

DGCA public statistics (domestic 2015 onward, airlines 2009 onward, international 2015 onward, monthly traffic reports 2014 onward) and the models in this report. Every number comes from a query you can open. It sees demand, not fares or schedules.

Source: DGCA (provisional). Answers describe demand structure; they never judge viability or profitability. Conversations are logged to improve the analyst.

هل أنتم مستعدون لبدء الحوار؟

اكتشفوا كيف تبني ArkOne هياكل الحوكمة وإدارة البرامج لتحقيق عوائد قابلة للقياس من الذكاء الاصطناعي.

احجز مكالمة استكشافية