Dev Diary: Rewiring the XM Network Under Extreme Pressure

Previously, I noted here that we migrated Ingress to a new database back in Mar 2026. The migration looked good initially. For two months, across multiple playtests, +Gamma Hyderabad, +Gamma Buenos Aires, two First Saturdays, and the Orion Global Op, we saw everyday gameplay latency improve and most game actions felt smoother.

But sometimes compounding factors can happen behind the scenes before reaching a tipping point, followed by an abrupt and unexpected collapse. Gradually, then suddenly. On May 16, the spike in Agent activity exposed critical bottlenecks that overwhelmed the new database. Auto-scalers were too slow to scale up. Game action logs turned into logjams. Ghost records refused to die.

Following Orion Prague, our team declared an emergency code red to completely rewire the Ingress backend in 13 days. This meant working nights and weekends through the Independence Day holiday to “move fast and don’t break things” before Orion Kure and Orion Jersey City. If we failed to ship the following fixes in time, then we knew Agents everywhere would experience general Scanner app instability and login issues. Again. (No pressure.)

Preheat the Anomaly Oven

As part of our preparation for each live event, we scale up the Friday before every Anomaly weekend. During Orion Prague, the database still took 40 minutes to analyze the traffic spike and move the right data to continue to serve requests. In addition to provisioning more machines, we now also specify exactly what data will be heavily requested before the first Anomaly measurement starts.

Remove Logjams

We used to maintain a separate database table of most Agent actions, indexed by timestamp. This was great for querying historical data, like generating a list of actions that happened between 14:00 and 14:15 for example. However, this was terrible for live write performance.

Since every action during the Anomaly happened right here, right now, every log entry was fighting to jam itself onto the exact same endpoint of the index at the exact same time. It was impossible for the database to keep up. We removed this index, and our internal administrative queries are slightly slower as a result, but they are good enough for NIA work.

Bust The Ghosts

Within the database, we also maintain an index of Portals sorted by physical location and last update time. In everyday gameplay conditions, this index is highly efficient. The Scanner app can query what changed nearby since it last checked and get a near-instant answer. But in Anomaly conditions, lots of Portals are updated lots of times.

Each time a Portal’s state was updated, the database created a new entry and added it to the top of the list, leaving behind a ghost entry that was marked for deletion later. There is a minimum cleanup time of one hour, but the spike in Agent activity generated 100X more ghost entries than real ones, which choked the system.

Our solution was to remove the last update time from the index ordering. Now, the database scans all nearby Portals and sorts them on-the-fly. Similar to how removing the logjam earlier meant slightly slower admin queries, there is a tradeoff here, too. By skipping ghost entries altogether, this brute-force method is now significantly faster in Anomaly conditions but slightly slower (less than 1.0ms) in everyday gameplay conditions.

Reduce the Blast Radius

Suffice to say, we did not want to risk repeating what happened on May 16, so anything and everything that might help mitigate these issues was considered. We know last minute changes (Calvinball) disrupt months of strategic planning that go into an Anomaly, but we also felt that we needed the extra headroom that could come from reducing the number of Portal updates at the edges.

We reduced XMP attack radius so that fewer Portals would be damaged by each XMP fire, which in turn also reduced the total number of server updates required. In order to balance that reduction, we increased XMP attack strength within that smaller attack radius. Fewer Portals getting attacked by a single Level 8 XMP meant fewer database calculations per Agent action, multiplied across hundreds of Agents firing thousands of XMPs in an Anomaly.

Move to In-Memory Updates

Finally, the Big Kahuna. On our previous database, Anomaly gameplay was by no means perfect, but the new database could not keep up. The biggest challenge we faced was how the new database handled rapid-fire updates to the same Portal.

Solving this required a seismic shift in how Ingress manages updates to the Portal Network. Instead of writing every update to the database instantly, we moved to an in-memory batching system that allows game actions to process in faster, temporary server memory first.

To recap, in 13 days our tiny but mighty Ingress team:

  • Updated over 200 files.
  • Added over 22,000 lines of code.
  • Removed over 5,000 lines of code.
  • Drank over 200 cups of coffee.

The Results

On May 30, at Orion Kure and Orion Jersey City, the new server architecture held up under extreme pressure. We avoided another weekend of Scanner app instability and global login issues, and our monitoring systems reported the lowest Anomaly gameplay latency on record.

Comparing Nov 15 2025 +Beta Taoyuan and +Beta The Hague, which ran on the old database, to May 30 Orion Kure and Orion Jersey City, which were our first Anomalies to run with pre-warmed Anomaly Sites, updated tables, the temporary reduced XMP range, and in-memory Portals:

  • Before (Nov 15 2025): 99th percentile latency peaked at 9X over 2025 baseline latency.
  • After (May 30 2026): 99th percentile latency peaked at 2X over 2025 baseline latency.

And we were not done yet. Our goal was to unblock Anomalies so they did not take all of Ingress down, and then continue to improve both the Anomaly and the everyday gameplay experience over time. Days before Orion Geneva and Orion Lima, we pushed fixes for Portals getting stuck that required a Like or thumbs up on the Portal photo to get un-stuck, and for issues that blocked some Agents from completing missions on Mission Day.

Orion Kure, Orion Jersey City, Orion Geneva, Orion Lima, Apollo Helsinki, and Apollo Bogotá have all run on the new in-memory system. However, we still have significant work ahead to be able to switch to in-memory Portals outside of Anomaly weekends.

Next steps

As some Agents have speculated, we use the new server architecture during Anomaly weekends and then disable it the following Monday. With each Anomaly, we collect more data, and Agents help us identify new issues to investigate further. And so it goes.

The new system can hold a Portal’s state in memory, gather gameplay actions as they happen, and then write the accumulated results to the database in batches. Portals, Links, Fields, Shards, and even dropped items all share the same table, so every game action in Ingress had to be rewired and it had to be done in a very short amount of time. If we missed even one action, the results would be catastrophic.

We need more time, but we are confident we will finish this work and provide the long-term stability and speed that the XM Network requires. Thank you for your patience, your honest feedback and bug reports (especially ones with detailed steps on how to reproduce what you are seeing), and your continued dedication to Ingress. We will celebrate 14 years this Nov, and we plan to celebrate many more milestones to come.

—Brian, on behalf of the Ingress team


Note: The following was translated using machine translation (please pardon any unintentional mistakes):

以前、2026年3月にIngressを新しいデータベースへ移行したことをここでお伝えしました。当初、移行は順調に見えました。その後2ヶ月間、複数のプレイテスト、「Gamma Hyderabad」、「Gamma Buenos Aires」、2回の「First Saturday」、そして「Orion Global Op」を通じて、日常的なゲームプレイのレイテンシ(遅延)は改善し、ほとんどのゲーム操作がよりスムーズに感じられるようになりました。

しかし、ある限界点(ティッピング・ポイント)に達する前に、水面下で複数の要因が重なり合い、突然予期せぬ破綻を招くことがあります。「徐々に、そして突然に」事態は悪化するのです。5月16日、エージェントの活動が急増したことで、新しいデータベースの処理能力を超える深刻なボトルネックが露呈しました。オートスケーラーによる拡張(スケールアップ)の反応は遅すぎました。ゲーム操作のログ処理は滞留(ログジャム)を起こし、削除されるべき「ゴーストレコード」が残り続ける事態となりました。

「Orion Prague」終了後、チームは緊急事態(コード・レッド)を宣言し、13日間でIngressのバックエンドを全面的に作り直すことにしました。「Orion Kure」や「Orion Jersey City」の開催を控え、「迅速に動く(Move Fast)」と同時に「システムを壊さない(Don’t Break Things)」ことを両立させるため、夜間や週末、さらには独立記念日の祝日も返上して作業に取り組みました。もし必要な修正を期限内にリリースできなければ、世界中のエージェントが再び、スキャナーアプリの全体的な不安定さやログインの問題に直面することになるのは明らかでした。(プレッシャーは相当なものでした。)

アノマリーに向けた「オーブン」の予熱
各ライブイベントの準備の一環として、私たちはアノマリー開催週末の直前の金曜日にシステムをスケールアップ(増強)しています。「Orion Prague」の際、トラフィックの急増を分析し、リクエストを処理し続けるために適切なデータを移動させるのに、データベースは40分もの時間を要していました。そこで現在は、サーバーの増設に加え、最初のアノマリー計測が始まる前に、どのデータへのリクエストが集中するかを事前に特定し、対策を講じるようにしています。

ログ処理の滞留(ログジャム)を解消する
以前は、エージェントの操作の大部分を記録するデータベーステーブルを別途用意し、タイムスタンプでインデックス(索引)を付けて管理していました。これは、例えば14:00から14:15の間に発生した操作のリストを作成するなど、過去のデータを検索・抽出する際には非常に有効でした。しかし、リアルタイムでの書き込みパフォーマンスという点では、極めて非効率な仕組みでした。

アノマリー中のすべての操作は「今、この瞬間」に発生するため、すべてのログエントリーが、インデックスの全く同じエンドポイント(書き込み位置)に、全く同じタイミングで書き込もうと競合してしまっていたのです。これではデータベースの処理が追いつくはずもありません。そこで私たちはこのインデックスを削除しました。その結果、内部管理用のクエリ(検索処理)は多少遅くなりましたが、NIAの業務を行う上では十分な速度を維持できています。 「ゴースト」の排除
データベース内では、物理的な位置と最終更新時刻でソートされたポータルのインデックスを管理しています。通常のプレイ状況では、このインデックスは非常に効率的に機能します。スキャナーアプリは、前回の確認以降に近隣で何が変更されたかを問い合わせ、ほぼ瞬時に回答を得ることができます。しかし、アノマリー(大規模イベント)の状況下では、多数のポータルが頻繁に更新されます。

ポータルの状態が更新されるたびに、データベースは新しいエントリを作成してリストの先頭に追加し、後で削除される予定の「ゴーストエントリ」を残していました。クリーンアップには最低1時間の待機時間が設けられていますが、エージェントの活動が急増したことで、実際のエントリの100倍ものゴーストエントリが生成され、システムがパンク状態に陥りました。

そこで私たちは、インデックスの並べ替え条件から「最終更新時刻」を除外するという解決策をとりました。これにより、データベースは近隣のすべてのポータルをスキャンし、その場で(オンザフライで)ソートを行うようになりました。以前、処理の滞りを解消した際に管理用クエリの速度がわずかに低下したのと同様に、ここでもトレードオフが生じます。ゴーストエントリを完全に無視するこの力技(ブルートフォース)的な手法は、アノマリーのような状況下では大幅に高速化しますが、通常のプレイ状況ではわずかに(1.0ミリ秒未満)低速になります。

爆発範囲(ブラスト・ラジアス)の縮小
言うまでもなく、5月16日に起きた事態を繰り返すリスクは冒したくなかったため、これらの問題を軽減するあらゆる可能性を検討しました。直前の変更(いわゆる「カルビンボール」的なルール変更)が、アノマリーに向けて数ヶ月かけて練られた戦略的計画を台無しにすることは理解していますが、同時に、周辺領域でのポータル更新数を減らすことで得られる「余力」が必要だと感じていました。

私たちはXMPの攻撃範囲を縮小し、1回のXMP発射でダメージを受けるポータルの数を減らしました。これにより、必要なサーバー更新の総数も削減されました。その削減のバランスを取るために、縮小された攻撃範囲内でのXMPの攻撃力を高めました。レベル8のXMP1発で攻撃されるポータルが減るということは、エージェントの1アクションあたりのデータベース計算量が減ることを意味します。アノマリーでは数百人のエージェントが数千発のXMPを発射するため、この差は大きな影響をもたらします。

インメモリ更新への移行
最後に、最大の課題について触れます。以前のデータベースでもアノマリー中の動作は決して完璧とは言えませんでしたが、新しいデータベースはさらに対応しきれない状態でした。私たちが直面した最大の課題は、同一ポータルに対する連続的な更新(ラピッドファイア)を新しいデータベースがどのように処理するか、という点でした。この問題を解決するには、Ingressにおけるポータルネットワークの更新管理のあり方を根本から変える必要がありました。すべての更新を即座にデータベースへ書き込むのではなく、ゲーム内のアクションをまず高速なサーバーの一時メモリ上で処理する「インメモリ・バッチ処理システム」へと移行したのです。

まとめると、わずか13日間で、Ingressの少数精鋭チームは以下のことを成し遂げました。

200以上のファイルを更新。
22,000行以上のコードを追加。
5,000行以上のコードを削除。
200杯以上のコーヒーを消費。

結果
5月30日、「Orion Kure」および「Orion Jersey City」において、新しいサーバーアーキテクチャは極めて高い負荷に耐え抜きました。Scannerアプリの不安定な動作や世界規模のログイン障害といった事態を回避できただけでなく、監視システムは「Anomaly(アノマリー)」のゲームプレイにおけるレイテンシ(遅延)が過去最低レベルであったことを記録しました。

旧データベースで運用された2025年11月15日の「+Beta Taoyuan」および「+Beta The Hague」と、事前準備されたAnomalyサイト、更新されたデータベーステーブル、一時的に縮小されたXMP射程、そしてインメモリ・ポータルを採用した初のAnomalyである5月30日の「Orion Kure」および「Orion Jersey City」を比較すると、以下の通りです。

以前(2025年11月15日):99パーセンタイル値のレイテンシは、2025年の基準値の9倍に達しました。
その後(2026年5月30日):99パーセンタイル値のレイテンシは、2025年の基準値の2倍に留まりました。
しかし、私たちの取り組みはこれで終わりではありません。目標は、AnomalyがIngress全体を停止させるような事態を防ぐこと、そして時間をかけてAnomalyと日常のゲームプレイ体験の双方を向上させ続けることでした。「Orion Geneva」および「Orion Lima」の開催数日前には、ポータルの写真に「いいね(サムアップ)」をしないと状態が解消されない不具合や、一部のエージェントが「Mission Day」でミッションを完了できない問題に対する修正を適用しました。

「Orion Kure」「Orion Jersey City」「Orion Geneva」「Orion Lima」「Apollo Helsinki」「Apollo Bogotá」は、新しいインメモリ・システム上で運用されました。しかし、Anomaly開催期間以外の通常時にもインメモリ・ポータルへ移行するには、まだ多くの課題が残されています。

今後の予定
一部のエージェントが推測されている通り、私たちは新しいサーバーアーキテクチャをAnomaly開催期間中に使用し、その翌週の月曜日には無効化しています。Anomalyが開催されるたびにデータを収集し、エージェントの皆様の協力によってさらなる調査が必要な新たな問題が特定されます。こうしたプロセスを繰り返しています。

この新しいシステムは、ポータルの状態をメモリ上に保持し、リアルタイムで発生するゲームプレイのアクションを収集した上で、蓄積された結果をまとめて(バッチ処理で)データベースに書き込む仕組みになっています。ポータル、リンク、フィールド、シャード、さらにはドロップされたアイテムに至るまで、すべてが同一のテーブルを共有しています。そのため、Ingressにおけるあらゆるゲーム上のアクションの仕組みを根本から作り直す必要があり、しかもそれを極めて短期間で行わなければなりませんでした。もし一つでも見落としがあれば、壊滅的な事態を招きかねない状況でした。

作業完了までにはもう少し時間が必要ですが、私たちは必ずこの作業をやり遂げ、XMネットワークに求められる長期的な安定性と高速な動作を実現できると確信しています。皆様の忍耐強いお待ちの姿勢、率直なフィードバックやバグ報告(特に、問題の再現手順が詳細に記されたもの)、そしてIngressへの変わらぬご尽力に心より感謝申し上げます。Ingressは今年11月で14周年を迎えますが、今後も数多くの節目を皆様と共に祝っていきたいと考えています。

—Brian(Ingressチームを代表して)

Share this article