資深遊戲製作人Tom Ellis在社媒發帖,詳細說明了《魔獸世界:永恆》Beta測試上線時遇到的問題。原定計劃,在現場問答結束後,無限服就應開放,但技術故障導致了延期,其後又出現掉線、排隊和高延遲情況。

原貼:
Alright, now the dust has settled and you're just happily playing away, let's answer the most common question I've seen. 1. AHMAGERD, this game is 20 years old, how do you still suck at this!?! a. Heh, fair enough, sort answer is, this is our BETA PTR environment stack, it's not a Production stack by a long, long, long, long, long, long shot. Betas for WoW aren't usually huge meter movers so this environment just isn't that powerful. This led to the first problem you saw today, Shortly after we went live and when you got through the login queue (which is just http://B.Net protecting itself, rate limiting logins, normal!), you were disconnected. Ah. That's not quite right. This one took the longest of the issue today to figure out (and even then, was around half an hour I think?). Our little BETA environment only uses one REGIONAL realm. Hrm, we should do a panel on WoW infra one day so any of this makes sense. Anyway... This turned out to be a problem, but it took a while to find because all our services CPU and Memory were A-OK, bored even. And when the http://Battle.Net Game Service team took a look at their services that we were talking to and getting timeouts from, they were bored too. What the hell. Eventually engineers figured out that we were hitting the BGS metering system that protects their services from falling over under extreme load. Our tiny little REGIONAL with its two connections that were trying to login A LOT of users, even though it was no where near capped on CPU was completely screwing the math that system was using. We scaled that 2 to an 8 and immediately the flood gates opened once again. We all recognized we need more logging and notifications when this system is kicking in, but this is the first time anyone's seen it kick in when their wasn't an obvious CPU issue to go along with it. 20 years, always something new. What next... Ah yes, shortly after that, many of you saw long looting times, accepting quests etc, anything that interacted with the database. The PTR database that beta uses was taking a long time to respond to queries. Our intrepid Oracle Database Engineers had to perform a manual analyze on all the tables and setup some automated jobs to periodically keep doing it, this is because some of these tables were new and existing tables were seeing a crapton of inserts and new query behavior so everything just needed to be checked more often while the data was still changing so much. As soon as these were performed DB performance cleared up instantly. For any DB engineers, yes we say analyze because we're old, we know its DBMS_STATS. Restarts! We kicked you all off. A few hours in, we noticed our WORLD pools that run, well, the game simulation you play on, were running crazy hot on CPU and almost out of memory. We then started seeing the first VMs get hit by the OOM'killer as they threatened to take down the hypervisor. Classic engineers were able to identify that we were not correctly shutting down empty maps, so we were slowly, well, not that slowly, gobbling up CPU and memory. One quick fix through passed through QA but needs restarts to pick it up. We also took the opportunity to add some additional WORLDs just in case, and gave that REGIONAL service a extra friend, just in case. Marked live once again and you all poured in without fuss. Right now things are incredibly smooth for the first few hours of a public beta, Classic teams watching and fixing anything blocking that comes up but we're otherwise pretty much done for the night, enjoy! https://twitter.com/FwoiblesWoW/status/2100751737627779238
「這遊戲都20年了,為什麽你們技術還這麼爛?有道理,簡單來說,是我們測試服的環境架構不同,跟正式版不是一個量級,差了十萬八千里。B測通常沒有太多負載壓力,所以我們的配置不高,這也導致了今天的第一個問題。」
「剛上線不久,好不容易通過了登陸驗證,卻遇掉線,這是所有問題中最難查的一個,找原因就花了大概半小時。我們這個小小的B測環境只用到一個區域級伺服器。也許哪天我們該辦個講座,聊聊《魔獸世界》的基礎架構,這樣大家才能了解這些技術細節。」
「的確是個問題,排查花了不少時間。當時我們伺服器的CPU和記憶體占用率都是正常,甚至可說閒得發慌,在出現各種連接超時,大家都很納悶。不過,最終問題被發現,原因是觸發了BGS的流量計量系統,雖然CPU負載遠未達到上限,但僅憑兩條連接試圖處理海量用戶登錄請求的行為,就徹底打亂了該系統所依據的計算邏輯。於是,我們將連接數從2調到8,流量閘門瞬間重新打開。我們都意識到,系統介入需要更完善的日誌記錄和通知機制。這是大家第一次見到它在沒有明顯CPU負載問題的情況下觸發,幹了20年,什麽怪事都會發生。」
「隨後是物品拾取、接任務等涉及資料庫交互的操作延遲極高的問題。當時B測所用的PTR資料庫對查詢請求的響應非常慢。我們的Oracle資料庫工程師不得不手動分析各種表,並設置自動任務定期執行此操作。因為一些表是新建的,現有的表也面臨海量插入操作和新的查詢模式。因此在數據劇烈變動期間,必須更頻繁地進行檢查。這些措施一經實施,資料庫性能便立即恢復正常。」
「我們重啟了伺服器,把大家都踢下線。但在運行了幾個小時後,我們發現負責處理遊戲模擬邏輯的世界服務CPU負載極高,且記憶體即將耗盡。隨後,首批虛擬機開始觸發OOM Killer(記憶體溢出殺手進程)機制,甚至威脅到了宿主機的穩定性。經工程師排查發現,我們未能正確關閉空閒地圖,導致系統不斷吞噬CPU和記憶體,且速度非常之快。一個快速修複方案通過了QA測試,但需重啟服務才能生效。我們也藉此增設了一些「世界服務」以備不時之需,並給那個「區域服務」增加了一個「夥伴」作為冗餘備份。我們再次開啟了實時測試,大家也都井然有序地湧入其中。懷舊服團隊正密切關注並修復任何出現的阻礙性問題,除此之外,我們今晚的工作已基本收尾——祝大家玩得開心!」






