<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Calvin&#39;s Marbles</title>
  
  
  <link href="http://www.calvinneo.com/atom.xml" rel="self"/>
  
  <link href="http://www.calvinneo.com/"/>
  <updated>2026-06-08T11:04:20.281Z</updated>
  <id>http://www.calvinneo.com/</id>
  
  <author>
    <name>Calvin Neo</name>
    
  </author>
  
  <generator uri="https://hexo.io/">Hexo</generator>
  
  <entry>
    <title>新加坡之旅</title>
    <link href="http://www.calvinneo.com/2026/06/03/meet-in-singapore/"/>
    <id>http://www.calvinneo.com/2026/06/03/meet-in-singapore/</id>
    <published>2026-06-03T11:20:33.000Z</published>
    <updated>2026-06-08T11:04:20.281Z</updated>
    
    <content type="html"><![CDATA[<p>和老妈逛了逛新加坡。</p><a id="more"></a><h1 id="D0"><a href="#D0" class="headerlink" title="D0"></a>D0</h1><p>出于省钱的考虑，我们买了 PVG 往返 SIN 的飞机票。周六早上 9.35 起飞，所以我们周五就住在浦东机场了。PVG 的汉庭挺旧的，索性还算干净。</p><h1 id="D1"><a href="#D1" class="headerlink" title="D1"></a>D1</h1><p>早上六点就从酒店过去了，结果被骗了，短信说是 H 岛，实际上是 T2 的廊桥，所以我们在 PVG 等了好半天。幸亏飞机没有延误，所以我们接近是满血登上飞机。这个飞机还不错，是个 787，不仅有午饭吃，甚至每个人还有 20 min 的免费 wifi 可以用。不过据说外航现在都接入马斯克的 starlink，全程有免费 wifi 了。anyway，我还是在飞机上看完了白日梦想家这个电影，然后我觉得我们冰岛没去白日梦想家小镇那也太对了，感觉实在没啥可看的，而且这个电影没有花多少镜头在这个小镇上。</p><p>下了飞机，我还处于没有电话卡的状态呢。我之所以这么勇，是因为之前查过樟宜机场是全 wifi 覆盖的（后来发现大部分地铁站也有免费 wifi）。</p><p>我们入境也挺麻烦的，那个护照扫描一直要 rescan，而我妈好像弹出一个比较奇怪的错误提示。后面工作人员挺不耐烦告诉我们要先去入境登记，所以我就跑到另一边的几个 ipad 那边入境登记。反正那个系统贼麻烦，填了一堆东西，末了我发现好像可以两个人写在一个单子里面，就可以少填一份，不过也没办法了。反正最后还是跌跌爬爬入境了，然后就是弄电话卡，然后去摆渡车。这个摆渡车还分了两个队，我傻等了两分钟，愣是等到一辆我们的车开走，才知道这个不是我们的队。所以我索性直接出去等了，不过也没等多久，车就到了，我们上去，总算到了 T2，走 T2 上了地铁。</p><p>我们先去的酒店，走索美塞下车，走一会就到了。不过这里路口很扯，那个 Orchard 路，一边有红绿灯可以过，一边没有。去宾馆就还是挺顺利的，也没要交什么城市税，也没有刷信用卡授信。这个宾馆其实还行，无非就是床小了点，设备旧了点。</p><p>放下东西，喝了点厕所的水，我们就出门。zhangqi 还在外面，所以我先带我妈去逛逛附近的福康宁公园。</p><p>晚上和 zhangqi 吃了饭。好久不见，原来他已经工作了。据说是在 HW 新加坡 branch，然后他们那并不加班。</p><h1 id="D2"><a href="#D2" class="headerlink" title="D2"></a>D2</h1><p>今天一天都是植物之旅。</p><p>早上首先是按照我从小红书临时搜到的线路 city walk，第一站坐到 city hall，看一下几个教堂。应该是周日的缘故，所以大家还在做弥撒，教堂不太能走进去。后面看了下旧禧街警察局，发现不就是昨天从福康宁公园下来的地方么？然后仔细一看，这边好像是新加坡现在的什么文化和青年部。再往前走就又是 clay quay 了，我们就没有再进去逛，而是往牛车水方向走。不过我实际感觉这段又累又没啥可看的，有个楼上涂鸦了一个花旦，我妈站在天桥上拍了下。进入了牛车水附近，路牌也会标注中文，还是挺有意思的。总而言之，在经过了老巴刹之后，我们就开始往喷水狮子那边走。路上有趣的是我看到了头条的新加坡办公楼。</p><p>到了喷水狮子那里，遮蔽越来越少，原来还能有连廊或者树木遮阳，现在直接在海边暴晒了。我们拍了拍现在已经近在咫尺的金沙三炷香，然后去狮子那边。狮子也没啥了不得的，我们就几个角度拍了照，就准备走了。</p><p>我看了下去滨海湾的公交，最后还是决定走路过去。不过走路过去好像也不是很远。这里我原以为滨海湾花园就是在金沙那三炷香顶上，后来才知道，这是一个专门的地方，在金沙的东边。一路上走着，听见说滨海湾的花穹和云雾林都很冷，想着我今天短袖短裤别给搞感冒啊。滨海湾花园有点类似于一片大公园，我们需要走一个小匝道绕进去。进去之后啥都没有，我还是看了路牌，换了一对导航才确信自己走对了。不过深入走了一段，很快就能看到那些长得很神奇的人造树，我确定自己走对了地方。然后就是买票，这边如果现场买票贼贵，一个景点都要 50 新。所以我搜了下 klook，两个人才 240，给我便宜哭了。我甚至担心这个是假的，不过我们很顺利就进去了。</p><p>花穹进去的通道，觉得很冷，但是等到了那个冷室里面，反而不冷了。这是因为冷室是玻璃的，所以太阳照进来暖洋洋的。这个冷室挺大的，分为了好多区。我刚进去往左走，分别是南非和南美。然后再深入，则可以往下或者往左上走。往左上走是去加州区，往下走则可以到一个庭院，里面是土耳其主题的花卉展，感觉最近是在和土耳其搞活动。说到这个，其实刚才一进来就看到一个土耳其航空的飞机，这我总算是回过味来了。</p><p>回到出口那，其实右边还有一堆地方没探索，主要是猴面包树区和多肉植物区。比较有意思的是这个区域中还陈列了一堆插花作品，我拍了一些，感觉都挺有创意的。</p><p>从花穹出来，就直接去云雾林了，距离很近，其实这两个甚至共用了一个出口。一进云雾林，就看到一个巨大的瀑布，非常壮观。然后再往前走就是一些恐龙，比如霸王龙，梁龙等等，头还能动，还有各种声光特效。我感觉挺违和的。不过云雾林非常有意思的地方是你可以坐电梯到 6 楼，然后爬到 7 楼，再往下走。因为这个展区实际上是为了展示热带雨林从高海拔到低海拔的植被的变化，所以这么做我觉得挺有意思的。类似的，我们晚上在国家兰花园里面，也看到一个类似的展区，按照海拔展示了不同的兰花。</p><p>云雾林有个特色就是下雾，进门去的时候，我们留意了下，下一场是 14.00。我原以为他会喷出一种西游记那样的烟雨蒙蒙的仙境的感觉，但实际上它就是搞了点小雾气洒洒水而已。</p><p>中午在滨海湾的时候和 Bugen 约了晚饭，是 Little Elephant 啥的店，一个泰餐。我看了下，发现它离植物园也不远，仔细研究了下，我发现可以从植物园北门逛到南门，这样正好可以坐地铁去那边。</p><p>从滨海湾花园出来，我们又是绕来绕去兜了不少路，走到一个商场的 B2 去吃一个肉骨茶，据说是米其林，然后点评也是排名前几的。结果一到那边发现 B2 是个食阁，所以米其林这么没牌面的么？食阁大概就是没有一个专门的店面，只有一个档口，然后中间有一些共享吃饭的地方。把饭拿过来之后，就在那边吃，吃完要自己送餐盘到规定的地方，然后还要按照是否清真来分类放置。</p><p>吃完饭，我们准备走路去地铁，前往植物园，路上还发现这个地方居然有赌场。我之前从没见过赌场，发现好像进去居然要登记护照。并且这个赌场看起来很奢华，占了一层，金碧堂皇的。我们后面走外面走，也好像路过这个赌场开在外面的门，但是是常闭的，写着什么“尊上禁止进入”之类的话，感觉又客气又严肃。</p><p>下了植物园地铁站，出到地面就是植物园。这个植物园是免费的，只是里面的兰花园要收费。我们的时间不多了，所以只能简单逛逛。刚进门的地方有一些竹子和攀爬植物，然后绕来绕去走到了一个植物园边缘的道路上。植物园是半开放的，这个道路走着走着就能走到外面的公路上，自然这条路也没啥风景可以看。但是我很快在地图上看到有个叫芳香植物的地方，所以我就想过去闻闻香味。所以我们又找了个岔路绕了点路回去，走到了一个小山上。在快走到的时候，我发现还有个园区是有关药用植物的，然后旁边还有个关闭着的园区写着是有毒植物，我们远远看了看，基本都不认识。不过药用植物里面有不少也挺香的，反而是走到芳香植物那个片区，我预想的那种芳香满庭的场景并没有发生，我们还是得凑到花上去闻才行。</p><p>从芳香植物下来之后，我们就准备前往兰花园了。不过路上我看到一个热带雨林片区，里面黑黢黢的，一开始不敢走，怕蚊子。不过后来看到两个老外进去了，所以我也返回，也带着老妈扎了进去。因为走的快，所以其实还好，没有被怎么咬到。中间还看到一棵大树，STRANGLING FIG，它的牌子上写着没有游客能够走过这个树而不驻足观看的。因为他真的特别特别粗。</p><p>到了兰花园门口，准备买票。我看了下 klook 好像没优惠，就直接走信用卡了。但好像其实你点进去，是有折扣的，不知道这个软件是怎么设计的。这个地方收费也很离谱，最便宜的只要一块钱，我们这种啥弱势群体/公民身份都不沾的就得 15 新奉上。入口的人还问我妈有没有满 60 岁，我也比较好奇这个是不是纯看诚信？</p><p>后来我总结，兰花主要是根和花。它的根有点类似于水仙花那样长，并且很容易寄生在别的树木上面，在一些地方，我看到它甚至是在钢筋上面的。</p><p>晚饭时间，点了冬阴功，感觉更酸辣。点了点香茅饮料水，感觉是真的挺好喝的，我和我妈后面又补了一杯。另外，沙拉也挺好吃的。我们在那里谈了不少东西，感觉非常传奇。有一点我比较震撼的是知道现在做 LLM 的人基本 90% 都是中国人，不过处于合规的原因，很多都得出国了。我还请教了 vLLM 和 SGLang 的区别，据说现在两个同质化很严重。SGLang 只是一开始引入的结构化生成语言（也就是得名）比较有意思，不过现在都是 OpenAI 的接口了。说到这，他也提到他现在在做的一些工作，我也对 vLLM 的接入层进行了学习了解。</p><p>然后我们就回宾馆了。</p><h1 id="D3"><a href="#D3" class="headerlink" title="D3"></a>D3</h1><p>今天一天都是动物之旅。</p><p>早上起的比较早，先去地铁站附近的大南洋吃了个早饭。是海鲜面啥的，感觉还行吧。</p><p>然后就是去动物园了，我们走索美塞上车，直接可以坐到卡迪。走卡迪下车，没几步就是 M2 的公交车了。其实在地铁站里面是有个路牌的，我当时还看着 google 地图找了一阵子。</p><p>这个 M2 也是奇怪，前面的车刚坐满，一个人没往上站，就让我们坐下一辆。但是下一辆又是站满了人，不知道是什么目的。总之也是晃晃悠悠才到了动物园片区。这个片区非常大，并不只是我们知道的五大动物园，还有一堆探险园啥的，都挤在一起。</p><p>我的计划是新加坡动物园和河川馆，因为飞禽世界在另一个区，而 night safari 大家都说啥都看不到。</p><p>先去了新加坡动物园，在 klook 上买的票，也是很省钱的，两个人才 358。不过这里坑的是最好买的时候就选一下日期，不然就要在一个很难用的网页上选日期，我们在门口搞这个卡了很久。另一个是最好每个人分开来买，这样优惠最高。我看了一些七拼八凑的攻略，也试了下旅途随身听，感觉实际都不如 Mandai 那个 App 有用。</p><p>我们墨迹到快十点才进园，然后我就直奔那个海狮表演。进去的时候，走了下树冠小路，比自己预期差了不少，因为只有一点点短。从上面下来之后，都是有凉棚的路，当然平行的也有一条树林小径，这个后面我们也走了一遍。</p><p>然后我们就走到动物园的中心，也就是餐厅，什么阿明餐厅啥的。在那边往左拐，再往右拐，走一会，就到那个剧场了。后来我发现这个海狮表演应该是这个动物园完成度最高的表演了。</p><p>海狮表演结束之后，我们连合照都没合照，就赶快去那个萌宠表演。这可要了大命，萌宠表演在另一个非常远的地方。我们后来回想，其实不如就在这里拍下照，等 11 点的另一场表演结束算了，萌宠表演可以下午看。而我们实际上是先急行军过去看了萌宠表演，然后下午两点半临走前，又绕回来看了这个剧场的另一个表演。</p><p>Anyway，我们走了点回头路，然后又穿过了整个绕来绕去的爬行馆片区，然后走到了一片 kid’s world 区域。我们在那里来回绕了半天，从冲淋区走到萌宠区，终于在一个类似于露天广场的地方找到了这个表演。当时表演已经进行了一部分了，我们看的部分是一只羊、一只猫、一只马和一只狗，我的第一反应是方舟动物园里面的萌宠不会就是照着新加坡动物园做的吧。不过这个表演与其说是才能展示，不如说是行为展示。羊象征性走了下秋千，但这个已经是最给面子的了。猫就是慢慢的走，然后穿过一个比一个小的圈圈。为了凑时间，甚至还邀请了一个大人和一个小孩上去钻更大的呼啦圈。马的展示就是看了四个刷子，然后刷一下马。因为马脸比较娇嫩，所以有个专门的刷子，感觉过的比人还好。那个叫 Lucas 的狗是最离谱的，啥都不做，就只会握手。</p><p>总之看完这个表演，我们才开始逛动物园。首先是先把萌宠这个区给逛完，然而我们也失算了，因为这个区实在是太乱了。我的想法是从 kids world 开始逆时针逛，所以大概是把来的时候的爬行馆给逛了。我觉得这是对的选择，因为新加坡动物园的爬行区确实很赞。南京红山森林动物园其实只是学了新加坡动物园的一些管理思想吧，比如不训练动物表演，但其实展出的动物类型还是差别蛮大的。比如我们就看到一个大蜥蜴，已经觉得很霸气了，但一看，是什么科莫多巨蜥，我一直以为这个是某个已经灭绝的牛逼生物呢。过了巨蜥，还看到不少超大的乌龟。而爬行馆里面更是五花八门，比如说有两个头的蜥蜴还是蛇，有传说中的西部菱斑响尾蛇。这里还有一个鸟/蝴蝶馆和脆弱雨林展区，我们也在稍后逛过了。鸟馆里面，与其说是鸟，我见到最多的是果蝠。爬到一个台子上面，可以离他们很近观看。雨林展区里面也是挺大的，感觉和爬行馆里面展示的东西差不多。</p><p>新加坡动物园还有一个特色的猴子比较多。从鸟馆出来之后，我也看到黑冠猴啥的。路上还看到有小蜥蜴在爬，还拍了视频。然后我们就离开这个片区往回走了，然后走到一个 Sensory Garden 里面，这里面是个南美巨骨鱼，那鱼还挺大的。然后这一圈又是一个猴子岛，里面有很多猴子，这一片其实就是灵长类王国。</p><p>出去之后，就是一个埃塞俄比亚大峡谷展区。里面展示了一些土著的建筑，以及一个非常大的狒狒园区。狒狒看起来还是非常离谱的，屁股又红又大。</p><p>最后一站看了亚洲象，然后就回到早上那个剧场看最后那个表演。中途，我准备抢 16.30 的票，发现网非常拉胯，硬是没抢到。</p><p>从新加坡动物园过来，就直接去河川馆了。这次吸取经验，让我妈自己买指定日的票了。进去之后，发现这个馆好像不是很大。我们是朝左走，好像是可以看大熊猫，虽然我也不是很喜欢大熊猫。它这个河川馆的意思是展示了全世界的各个河流的生态，比如尼罗河、澜沧江等等，基本上是一条单向的步道，走到大熊猫那边是长江展区。看完之后，相当于就走了一半了。跨过一个桥，就到了另一边。另一边主要就是 16.30 的表演，以及一个 5 新的船。我买了票，说是要排 30 min，实际上 12 min 就上了。我们一共坐了三次，3 4 5 排都体验过了。这个船还是挺好的，首先它不是很吓人，也没有溅很多水出来。然后，它路过的动物都是比较活跃的，比如食蚁兽、朱鹮、猎豹、卡皮巴拉，还有一些猴子等。</p><p>坐完船，其实前面就没啥了。有印象的主要是一个亚马逊主题，里面有一个非常深的水池，里面养了海牛等等之内的大动物。可以从一个海底隧道进去，然后绕着圈圈一直看到最上面。不过因为我去过大阪的海游馆，所以我觉得也就 soso。</p><p>从动物园出来，准备在卡迪地铁站附近的超市买点水果，都准备买了，发现原产地是澳大利亚，感觉没啥意思啊，就没买。这边水果品种也不多，4 个芒果 4 新好像倒是很便宜，但是特别大，感觉带不走。</p><p>因为 zhangqi 推荐了白鸡尾饭，所以今天我在 Orchard 站附近的一个商场的食阁里面找到了一家也是很有名的天天海南鸡饭。这个商场是真的破，感觉和兴化的招商城有的一拼。但里面剪头美甲真的好贵，剪头要 25 新，美甲要 100 新，而且也不是那种很装修豪华的店，就跟街边的小店一样，感觉新加坡是真的贵。天天家的点菜，我是真的不知道了，因为我不知道白鸡尾饭用英文怎么说。幸亏那个店员会中文，所以我跟他说他是听的懂得。好像这个就是海南鸡，只是部位不同。一般的是带骨头的，但是鸡尾是不带骨头的，并且还特别嫩。</p><p>从这个店出来，找到一家 fair price 超市，准备买点水果。这边不少水果奇形怪状，我找了半天，最后用 kimi 和 chatgpt 帮忙选了一个叫 salak 的水果。吃着挺好吃的，很香，肉也紧实，水不是很多，吃在嘴里还有点脆。</p><p>回到宾馆，问了下服务人员，说是可以免费寄存行李的，只要不过夜。</p><h1 id="D4"><a href="#D4" class="headerlink" title="D4"></a>D4</h1><p>今天起的更早了，原因是我想先出去逛一逛，然后回来洗个澡再退房，不然我一天汗流到屁股沟子里面很难受了。昨天晚上我就在纠结到底去哪里，早上去圣淘沙好像不太来得及，但是那些博物馆又是十点钟才开门。正好我妈昨天晚上说见了一大堆外国人，所以我寻思就带他去看看各个民族的街区吧。</p><p>所以第一站我们去了小印度。这时候，旅途随身听就起到了作用，对于这个景区，它规划还是挺好的，讲解也不错。我们先去了竹脚中心，我妈有点想吃印度飞饼，但是我们找了一个摊子，发现人家说没开门，也不知道真的假的，我看他倒是做的热火朝天的。anyway，我们逛一圈也没看到有啥可以吃的，就出去继续逛了。期间逛了一些某个华人的故居，甘地纪念馆，一些印度不同地方风格的寺庙。比较有趣的是在一条街上距离 30 米内有一家印度庙，一家观音庙和一家清真寺。</p><p>逛完小印度，准备去苏丹回教堂和甘榜格兰。不过发现与其坐地铁，好像走几步就到了。中间我还试了下这边的甘蔗汁机器，感觉不错。这边甘蔗比较有意思的，它好像一节节纺锤一样。走到武吉士附近，能看到一个 Rochr 运河，以及来福士医院。我妈觉得挺神奇的，为什么这个医院门口没有像江苏省人民医院门口全是人挤人。苏丹回教堂看着挺漂亮的，不过没开。往前走一点就到了甘榜格兰，它感觉就差强人意，其实就是个小庙。而且也没开放，还有墙围着，所以我们就准备往回走，去看看马来文化园。这个地方好像之前说黄循财来过，最近还在免费期。</p><p>到了文化园，还差几分钟 10 点，准备等一下。我妈上了个厕所，我找了半天入口，发现在另个建筑中。走进去发现这其实是一个小馆，并不是很大，两层楼，每层楼有四五个展厅。刚进去的时候，门开不下来，然后找到那边办公室，说是免费的，但是还是要先给我们一个 tag 贴在衣服上，比较有仪式感了属于是。一楼是新加坡国王，以及殖民相关的，比较有印象的是有个抓大象用的矛。二楼是近代相关的，还有一个专门的马来西亚电影艺术家展厅。这里比较有意思的是二楼的一个像船一样的东西，其实叫 Congkak。另外还有个刨椰子的东西。</p><p>我们是匆匆逛完这个文化园的，然后就快速往宾馆走。路上又看到那个回教堂，但是抓紧时间，也没有再进去，从外面看了下，里面全是人。到了宾馆，洗了个澡。我妈把我想扔掉的衣服又偷偷背回来了，不过我还是想办法扔掉了个白的。</p><p>从宾馆出来，这次就直接退房了。因为要去新加坡国家博物馆，所以路线还是和第一天一样，路过多美歌，以及福康宁树洞，这一次总算是可以看看长什么样子了。我们路过 Istana Park 的时候，我突然意识到，对面那个戒备庄严的地方，可能就是新加坡的总统府，后来发现好像确实是这样。</p><p>到了新加坡博物馆，发现太阳是真的毒辣。此时，刚好快 12 点，我寻思博物馆也不知道要逛多久，不如吃个东西。一搜，结果附近有一个便宜米其林，这就过去吃了。</p><p>吃完，就去博物馆。这个博物馆应该是我看过最垃圾的博物馆了。</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;和老妈逛了逛新加坡。&lt;/p&gt;</summary>
    
    
    
    
    <category term="游记" scheme="http://www.calvinneo.com/tags/游记/"/>
    
  </entry>
  
  <entry>
    <title>Vector Index 相关技术调研</title>
    <link href="http://www.calvinneo.com/2026/05/22/on-vector-index/"/>
    <id>http://www.calvinneo.com/2026/05/22/on-vector-index/</id>
    <published>2026-05-22T15:09:06.000Z</published>
    <updated>2026-05-22T09:18:38.027Z</updated>
    
    <content type="html"><![CDATA[<p>Vector Index 相关的一些总结吧。</p><a id="more"></a><h1 id="存储"><a href="#存储" class="headerlink" title="存储"></a>存储</h1><h2 id="综述"><a href="#综述" class="headerlink" title="综述"></a>综述</h2><p>目前存储格式基本上分两种思路：</p><ul><li>基于图的方法，如 HNSW、Vamana / DiskANN<ul><li>结构：向量是图里面的点。选择比较近的向量，或者能够向其他 cluster 连通的向量（Vamana）去连边。</li><li>搜索方式：从入口点开始，查找当前节点的所有邻居和待查询向量的距离，然后选择最近的跳过去。重复迭代，直到找不到更近的邻居。</li><li>特点：Accuracy 高，但如果图太大或者结构太复杂，IO 性能不是很好。特别地，HNSW 得全部放在内存里面。</li></ul></li><li>基于聚类的方法，如 IVF-PQ、SPANN、SPFresh<ul><li>结构：先用 KMeans 等算法对向量进行聚类，然后把这些向量都挂到对应的 cluster 下面，类似倒排索引一样（Inverted File / IVF）。</li><li>搜索方式：找到最近的几个 Centroids，把倒排索引里面的向量捞出来比对。</li></ul><ul><li>特点：对磁盘友好，数据局部性好。</li></ul></li></ul><h3 id="量化技术"><a href="#量化技术" class="headerlink" title="量化技术"></a>量化技术</h3><h2 id="Vamana-算法"><a href="#Vamana-算法" class="headerlink" title="Vamana 算法"></a>Vamana 算法</h2><p>Vamana 没有 HNSW 那么多层级，它只有一层图。这一层图中，既保留了 HNSW 中最下面的稠密层类似 kNN 那样的边，但是又保留了一些长程边。这样，兼顾了 HNSW 稀疏层的“快速跨越跨度”能力和稠密层的“局部精细搜索”能力。</p><p>因为只有一层，所以 Vamana 是把一个节点及其所有的邻居节点、连同它们的原始向量，物理上紧密地排列在磁盘上的相邻位置。当算法访问节点 $A$ 时，直接把 $A$ 的向量和它所有邻居的信息整块（Block）读入内存。Vamana 严格限制每个节点的最大邻居数（例如最多 64 个）。这样可以确保“节点+邻居信息”的大小绝对不会超过磁盘的一个 Page（通常是 4KB）。</p><h3 id="具体实现"><a href="#具体实现" class="headerlink" title="具体实现"></a>具体实现</h3><ol><li>搜索时的 Tradeoff：参数 $L$<ul><li>$L$ 设得大： 搜索时在内存里维护的候选队列更长，探索的分支更多。代价是更慢（计算距离和磁盘 I/O 次数增加），收益是更准确（Recall 更高，不容易漏掉正确答案）。</li><li>$L$ 设得小： 搜索时只看眼前最近的几个节点。代价是容易掉进局部最优解（准确率下降），收益是极速返回。</li></ul></li><li>建图时的 tradeoff：参数 $\alpha$<br> $\alpha$ 是 Vamana 独创的 RobustPrune 算法中的参数，通常 $\alpha \ge 1$。它的真正使命是控制长程边（跨区域高速公路）的生成，从而从根本上决定搜索的跳数和效率。<br> 假设在建图时，节点 $A$ 已经连了邻居 $B$，现在算法在评估要不要保留通向远处 $C$ 的边：<ul><li>当 $\alpha = 1$ 时（严格剪枝）： 算法非常“近视”。只要通过 $B$ 通过多条短边一点点挪去 $C$ 的距离在数学上还能接受，它就会把 $A$ 直连 $C$ 的边剪掉。最终图里全都是局部短边。由于缺乏跨越巨大空间的长程边，搜索时需要一步步缓慢挪动，跳数急剧增加，导致搜索非常慢。</li><li>当 $\alpha &gt; 1$ 时（例如 1.2，适度放松）： 算法变得更加包容，会刻意保留一些跨越较大空间的长程边。这正是 Vamana 能用单层图匹敌 HNSW 多层图的奥秘。这些长程边充当了“高速公路”，搜索时能极快地跨越广阔的向量空间逼近目标区域，不仅让搜索变快（跳数大幅减少），还能跳出局部死胡同（提高准确率）。</li><li>当 $\alpha$ 极大时（例如无穷大）： 算法完全不剪枝，直接退化成传统的 kNN Graph（每个点只连物理上绝对距离最近的几个点）。整张图彻底失去长程高速公路，且极易断连成一个个孤岛，导致全局搜索常常在中途断掉，准确率会暴跌。</li></ul></li></ol><h3 id="α-的判断条件"><a href="#α-的判断条件" class="headerlink" title="α 的判断条件"></a>α 的判断条件</h3><p>在 Vamana 为某个中心节点建图时，它会按距离从近到远遍历所有的候选节点，并使用以下公式来决定是保留还是剪除（Prune）一条边：<br>如果在候选池中，存在一个节点 $p’$，满足以下条件：</p><p>$$\alpha \cdot d(p^*, p’) \le d(p, p’)$$</p><p>那么，系统就会剪掉中心节点 $p$ 直连 $p’$ 的边（将其从候选池中剔除）。<br>这里面的三个变量代表什么？</p><ul><li>$p$：当前正在建边的中心节点。</li><li>$p^*$：刚才已经被 $p$ 选中的近邻节点（已经决定要连边了）。</li><li>$p’$：还在候选池子里，等待被评估要不要和 $p$ 连边的目标节点。</li><li>$d(x, y)$：表示两个节点之间的距离。</li></ul><p>不妨令 alpha 为 1，这也意味着：如果目标节点 $p’$ 距离已选邻居 $p^*$ 的距离，比它距离中心节点 $p$ 的距离要小，那么 $p’$ 就不要连一根线去 $p$ 了。那么搜索的路径就是</p><p>$$<br>p \rightarrow p^* \rightarrow p’<br>$$</p><h3 id="alpha-等于-1-或者无穷大的时候，都没有长程高速公路？"><a href="#alpha-等于-1-或者无穷大的时候，都没有长程高速公路？" class="headerlink" title="alpha 等于 1 或者无穷大的时候，都没有长程高速公路？"></a>alpha 等于 1 或者无穷大的时候，都没有长程高速公路？</h3><ul><li>当 $\alpha = 1$ 时（极端吝啬）<br>  高速公路死于：“过度追求最短路径，把长边当成了浪费”。<ul><li>逻辑： 当 $\alpha = 1$ 时，剪枝机制极其严格。只要 $A$ 觉得“通过 $B$ 走到 $C$ 的距离 $\le A$ 直达 $C$ 的距离”，$A$ 就会无情地把直达 $C$ 的长边剪掉。</li><li>结果： 在这种极度“抠门”的策略下，算法认为任何长跨度的边都是“冗余的、不划算的”，因为总能找到一系列小短边拼凑出一条路径。</li><li>资源状态： 节点手里的 64 个连边名额（$R$）可能都没用完。但它就是不愿意建长边，最后图里全是非常精简、互不重叠的小短边（在数学上这非常接近 Delaunay 三角网格）。</li></ul></li><li>$\alpha = \infty$ 时（极端短视）<br>  高速公路死于：“名额被短边全部占满，根本没有余力去连远方”。<ul><li>逻辑： 当 $\alpha = \infty$ 时，剪枝机制彻底关闭（不再判断抄近路的问题）。算法会严格按照绝对物理距离，连接周围的点。</li><li>结果： 因为距离中心节点最近的点，肯定是它身边的局部点。这些局部点会迅速消耗掉那 64 个连边名额（$R$）。</li><li>资源状态： 当算法扫描到远处的节点（潜在的高速公路目标）时，发现名额已经满了。</li></ul></li></ul>]]></content>
    
    
    <summary type="html">&lt;p&gt;Vector Index 相关的一些总结吧。&lt;/p&gt;</summary>
    
    
    
    
    <category term="数据库" scheme="http://www.calvinneo.com/tags/数据库/"/>
    
    <category term="数据结构" scheme="http://www.calvinneo.com/tags/数据结构/"/>
    
  </entry>
  
  <entry>
    <title>例论 TiFlash 的 HTAP 架构对排查不一致问题的影响</title>
    <link href="http://www.calvinneo.com/2026/05/06/tiflash-inconsistency-retro/"/>
    <id>http://www.calvinneo.com/2026/05/06/tiflash-inconsistency-retro/</id>
    <published>2026-05-06T15:09:06.000Z</published>
    <updated>2026-05-09T08:00:00.196Z</updated>
    
    <content type="html"><![CDATA[<p>以几年前的一个 case 为例，讨论了 TiFlash 的 HTAP 架构对排查不一致问题的影响，也讨论了如何用 AI 来减少这一类问题的调查时间。</p><a id="more"></a><h1 id="正文"><a href="#正文" class="headerlink" title="正文"></a>正文</h1><h2 id="结论"><a href="#结论" class="headerlink" title="结论"></a>结论</h2><p>这个 bug 的表象是：endless consistency 测试中，<code>insert</code> 负载下同一个 <code>start_ts</code> 对拍 TiKV 和 TiFlash，TiFlash 偶发少返回若干行，例如 <code>tikv: [35068980], tiflash: [35068960]</code>。再次用同一个 TSO 查询，结果会自动恢复一致，说明底层数据没有永久丢失，也不是被后续 snapshot 覆盖修好的持久化损坏。</p><p>最终根因不是 ReadIndex 没有推进 <code>max_ts</code>，也不是 TiKV memory lock、bypass lock、Region Split 本身。真正的问题在 DeltaMerge 的 <code>DeltaIndex</code> 复用：一个较早的读已经拿到了旧的 segment snapshot，但尚未完成 place index；另一个较晚的读先执行，把共享 <code>DeltaIndex</code> 推进到了更新的状态。这个更新后的 <code>DeltaIndex</code> 中如果包含“重复写入的同一批 row_id + version”，会把较早 snapshot 中仍应该可见的那批重复 tuple 标记为 deleted。较早的读继续执行时复用了这个过新的共享 <code>DeltaIndex</code>，于是漏读了这些行。</p><p>PR <a href="https://github.com/pingcap/tiflash/pull/9000" target="_blank" rel="noopener">#9000</a> 的修复就是围绕这个点：</p><ul><li>在 <code>DeltaTree</code> 里记录 <code>max_dup_tuple_id</code>，place insert 时记录发生重复 tuple 的最大 tuple id</li><li>clone <code>DeltaIndex</code> 时除了检查 delete range 数量不超过 snapshot，还要检查 snapshot 是否覆盖了共享 <code>DeltaIndex</code> 中所有 duplicated tuples<ul><li>如果不覆盖，就不能复用共享 <code>DeltaIndex</code>，必须为当前 snapshot 重新 place。</li></ul></li></ul><h2 id="最小化后的触发序列"><a href="#最小化后的触发序列" class="headerlink" title="最小化后的触发序列"></a>最小化后的触发序列</h2><p>issue 评论中 J-H 给出的抽象复现可以概括成：</p><ol><li>线程 A 通过 <code>resolveLocksAndReadRegionData</code> 写入一批行 <code>batch_1</code>，随后获取 segment snapshot <code>snap_a</code>。</li><li>线程 C 又写入一批和 <code>batch_1</code> 具有完全相同 row_id 和 version 的行 <code>batch_2</code>。</li><li>线程 B 获取更晚的 segment snapshot <code>snap_b</code>。</li><li>线程 B 先于线程 A 执行读流程，并更新 segment 的共享 <code>DeltaIndex</code>。</li><li>线程 B 更新 <code>DeltaIndex</code> 时，因为 row_id/version 重复，<code>batch_1</code> 在 delta tree 中被 delete 掉。</li><li>线程 A 继续使用 <code>snap_a</code> 读，但 clone/reuse 了线程 B 推进后的共享 <code>DeltaIndex</code>，导致 <code>batch_1</code> 对线程 A 不再可见，从而少读。</li></ol><p>这个错误是 transient 的：后续读通常会拿到更完整、更靠后的 snapshot，delta 覆盖了重复写入的全部 tuple，<code>DeltaIndex</code> 和 snapshot 再次匹配，所以结果自然恢复。</p><h2 id="现场现象和证据链"><a href="#现场现象和证据链" class="headerlink" title="现场现象和证据链"></a>现场现象和证据链</h2><p>初始 issue <a href="https://github.com/pingcap/tiflash/issues/8845" target="_blank" rel="noopener">#8845</a> 在 2024-03-14 提交，描述是 consistency test 中 <code>select count(*)</code> 对拍失败，TiFlash 少 20 行。Issue 被标记为 storage critical bug。</p><p>关键现象：</p><ul><li>workload 基本只有 insert，没有注入错误。</li><li>错误窗口很短，同一个事务或同一个 TSO 后续查询会收敛到正确结果。</li><li>多次现场都伴随 Region Split、Remote Read、Learner Read、lock resolve 或并发读写 DeltaMerge。</li><li>不是所有 TiFlash 节点都错，通常是某个节点的某次 TableScan 少了固定数量的行，例如 20、50、863。</li><li>检查同一时间之后的 apply snapshot，排除了“后续 snapshot 覆盖坏数据导致恢复”的解释。</li></ul><p>4 月下旬的日志把方向从事务层拉回到了 DeltaMerge：</p><ul><li>同一 region、同一 applied index 下，来自 Raft Apply 线程的写入和来自 Learner Read/resolve lock 路径的写入发生并发<br>  这其实是一个关键线索，因为它能够极大的缩小范围。实际上如果 C-N 当时没有看到这条线索，那么会在事务层上浪费更多的时间。<br>  C-N 当时提到，<strong>不妨假设事务层是有问题的</strong>，那么在这种情况下，Delta Index 也有问题。</li><li>同一批 50 条记录被两条路径写入。第一次读的 snapshot 中 <code>fully_indexed=false</code>，第二次读已经 <code>fully_indexed=true</code></li></ul><p>两次 snapshot 行数差和 delta index 行数差相差正好 50，吻合“重复 tuple 被过新的 delta index delete 掉”的模型。</p><h2 id="定位过程：难点主要在找到-DeltaIndex"><a href="#定位过程：难点主要在找到-DeltaIndex" class="headerlink" title="定位过程：难点主要在找到 DeltaIndex"></a>定位过程：难点主要在找到 DeltaIndex</h2><p>这个问题在定位到 DeltaIndex 之后，修复并不复杂，真正困难和耗时的是前面的定位：现象强烈指向事务/ReadIndex/lock/Remote Read，但最终根因藏在 DeltaMerge 的读索引复用。</p><p>早期首先确认了几个事实：</p><ul><li>TiFlash 少的是某一次 TableScan 的行数，不是聚合或 TiDB 层统计错误。把同一查询拆到各 TiFlash 节点看 <code>MPPTaskStatistics</code>，发现只有某个节点某次 scan 少了 20/50 行。</li><li>用相同 <code>start_ts</code> 后续再查会恢复，说明不是数据永久丢失，也不是 PageStorage/DeltaMerge 持久化写坏。</li><li>错误现场经常伴随 Region Split 后 epoch mismatch 触发 Remote Read，因此最早怀疑 remote read 的 key range、region split 处理、apply snapshot 覆盖等路径。</li><li>检查错误时间点之后的 apply snapshot，发现没有发生能解释“恢复正确”的 snapshot 覆盖行为，于是排除了 apply snapshot 覆盖坏数据。<br>  这个判断是很重要的，因为在检查 raft 路径导致的不一致问题的时候，需要首先区分是否是 apply snapshot 导致的。对应的后续调查方案是不一样的。</li></ul><p>接着排查转向事务可见性。C-N 在 issue 第一条评论里提出过一个很合理的怀疑：</p><ul><li><a href="https://github.com/pingcap/tiflash/pull/3971" target="_blank" rel="noopener">#3971</a> 引入了同 region 的历史 ReadIndex 结果复用</li><li>如果两个事务互不可见，<code>start_ts</code> 和 commit record 在 Raft log 中的顺序不一定一致，那么较小 TSO 复用较大 TSO 的 read index 可能读不到后面的 commit record<br>  这个假设解释力很强，因为现象正好是“TiFlash 少读已提交数据”。</li><li>但后续 J-H 指出，Async Commit 的 <code>max_ts</code> 推进理论上应该防止这个问题，所以它不能直接作为根因。</li></ul><p>之后又集中排查了 <code>max_ts</code>、memory lock 和 bypass lock：</p><ul><li>TiFlash 的 read index 路径一度被怀疑没有真正通过 TiKV 的 <code>ReadIndexObserver</code> 推进 <code>max_ts</code>，因为旧的 <code>read_index</code> RPC 已经不再使用，而 TiFlash 看起来发送的是 <code>RaftCmdRequest{cmd_type=ReadIndex}</code>。<br>  后来通过继续跟 TiFlash proxy 到 TiKV leader 的交互确认，TiFlash 虽然内部使用 <code>RaftCmdRequest</code>，但真正和 TiKV 交互仍走 <code>MsgReadIndex</code> raft message，路径是合理的。</li><li>使用 TiKV 侧临时日志确认：TiFlash Proxy 发出的 <code>MsgReadIndex</code> 带上了目标 <code>start_ts</code>；TiKV 收到后确实打印了 <code>advance max_ts to ...</code>，也返回了对应 read index。这排除了“完全没有推进 <code>max_ts</code>”的简单解释。</li><li>memory lock 也被排除<br>  因为：多个复现场景中错误附近没有足够的 memlock 事件；有些焦点 region 的 <code>bypass_locks</code> 为空；TiFlash 遇到 lock 后触发 Remote Read、resolve lock、重试的行为在日志上能解释，但不能解释最终少行。</li><li>bypass lock 也被排除<br>  因为：相关日志显示读事务并没有依赖异常的 bypass lock 路径，且 lock 的 <code>min_commit_ts</code>/<code>commit_ts</code> 关系不能完整解释“同一 TSO 后续恢复”的现象。</li></ul><p>这个阶段最迷惑的是“矛盾日志”：</p><ol><li>TiKV 已经 advance <code>max_ts</code> 并返回 read index，例如 region 459 返回到 313679。</li><li>但 TiFlash 随后在 apply 313680 时仍看到 <code>commit_ts &lt; read_tso</code> 的记录。</li></ol><p>这看起来像 follower read 与 Async Commit 的线性一致性问题。后续还去咨询 TiKV 侧，分析 async commit、read through lock、<code>min_commit_ts</code> 推进规则。这个方向耗时很大，因为它确实和一致性语义相关，而且日志表面上支持。</p><p>后续这个问题进入了僵局。最终 C-N 发现了下面的问题：</p><blockquote><p>转⽽考虑，⽆论是否承认 max_ts 有问题，⾄少确认已经 apply 到了 79 这个 raftlog，那么：</p><ol><li>要么能看到 lock</li><li>要么能看到 commit<br>不应该出现既没有看到 lock ⼜没有看到 commit 的情况。因此准备转⽽调查 resolvelock 和 raft command 两个地⽅的⾏转列的过程。</li></ol></blockquote><p>这是整个问题解决的关键，因为它揭示了在经过了那一团浆糊一样的事务问题之后，无论这时候的数据是否已经有问题，那么后续的处理也是有问题的。</p><p>C-N 把两次读同一个 segment 的信息做了对比：</p><ul><li>第一次读拿到的 snapshot：<code>fully_indexed=false</code>，snapshot 中的 stable/persisted/mem rows 总数较少。</li><li>第二次读拿到的 snapshot：<code>fully_indexed=true</code>，snapshot 行数更多，<code>DeltaIndex</code> 也更靠后。</li><li>两次 snapshot 行数差是 4860，两次 delta index placed rows 差是 4810，差值正好 50。</li><li>同时 Learner Read 线程和 Raft Apply 线程对同一 region、同一 applied index 写入了同一批 50 条记录，版本和 handles 完全吻合。</li></ul><p>这组证据把问题从“为什么读不到已经 commit 的 MVCC 数据”改写成“为什么旧 snapshot 会使用包含额外重复 tuple 删除信息的过新 DeltaIndex”。随后 JHL 在 PR #9000 中抽象出单测：</p><ul><li>线程 A 先拿旧 snapshot，线程 B 先读并推进 shared delta index</li><li>然后线程 A 复用这个 shared delta index 时少读。</li></ul><p>PR 评论里最小测试曾复现 <code>count1 = 118, count2 = 128</code>，这直接证明了问题在 DeltaIndex 和 snapshot 的匹配关系，而不是事务层。</p><h2 id="为什么这么难查"><a href="#为什么这么难查" class="headerlink" title="为什么这么难查"></a>为什么这么难查</h2><ol><li>症状像事务一致性问题，但根因在存储读索引复用。TiFlash 少读的行对应的 <code>commit_ts</code> 小于读 TSO，且现场经常出现 ReadIndex、Async Commit、memory lock、bypass lock、Remote Read、Region Split，因此很自然会先怀疑 <code>max_ts</code> 推进、lock resolve 或 ReadIndex 复用。</li><li>问题会自动恢复，容易误导排查。相同 TSO 之后再查是正确的，说明数据既没有永久丢也没有损坏；这排除了很多常规方向，但也让现场窗口很短，必须在错误发生时保留 segment snapshot、DeltaIndex、KVStore、日志和 region 信息。</li><li>复现条件苛刻。它需要高并发读、Remote Read/resolve lock 触发的写入、Raft Apply 写入、相同 row_id/version 的重复 tuple、旧 snapshot 和新共享 DeltaIndex 交错复用。早期本地集群和 IDC 集群很难稳定复现，后来通过调小 region 大小、提高 split 频率、打开 failpoint 增加 remote read 才提高概率。</li><li>中间出现了多个合理但非最终的假设。PR <a href="https://github.com/pingcap/tiflash/pull/3971" target="_blank" rel="noopener">#3971</a> 引入的 batch/async read-index 复用逻辑一度被怀疑；<a href="https://github.com/pingcap/tiflash/pull/8873" target="_blank" rel="noopener">#8873</a> 增强了 ReadIndexWorker 日志；<a href="https://github.com/pingcap/tiflash/pull/8874" target="_blank" rel="noopener">#8874</a> 增加写入 version 和 segment snapshot rows 日志；<a href="https://github.com/pingcap/tiflash/pull/8928" target="_blank" rel="noopener">#8928</a> 用来检查 async task 测试/日志路径。这些帮助排除方向，但不是最终修复。</li><li>TiKV/TiFlash 日志存在时钟漂移和跨组件链路。文档里曾用 TiKV 加日志确认 TiFlash Proxy 确实发了 <code>MsgReadIndex</code>，TiKV 也 advance <code>max_ts</code> 并返回 read index；但 TiFlash 随后仍看到看似 <code>commit_ts &lt; read_tso</code> 的写入。这让问题一度看起来像 Async Commit 与 follower read 的线性一致性问题。后续才确认这不是最终矛盾，而是并发写入和 DeltaIndex 复用造成的读视图错配。</li></ol><h2 id="从-TiFlash-架构看，为什么一致性问题更难查"><a href="#从-TiFlash-架构看，为什么一致性问题更难查" class="headerlink" title="从 TiFlash 架构看，为什么一致性问题更难查"></a>从 TiFlash 架构看，为什么一致性问题更难查</h2><p>TiFlash 的 HTAP 架构选择是“TiKV 行存主副本 + TiFlash 列存 learner 副本”。官方文档中也说明，TiFlash columnar replica 是通过 Raft Learner 异步复制得到的：<a href="https://docs.pingcap.com/tidb/stable/tiflash-overview/" target="_blank" rel="noopener">TiFlash Overview</a>；VLDB 论文 <a href="https://vldb.org/pvldb/vol13/p3072-huang.pdf" target="_blank" rel="noopener">TiDB: A Raft-based HTAP Database</a> 也描述了 TiFlash learner 接收 Raft log，并把 row-format tuple 转成 columnar representation。这个架构带来了很好的 HTAP 隔离和列存性能，但一致性排查天然更复杂。</p><ol><li>TiFlash 不是事务提交路径上的 voting replica，而是 learner columnar replica。写入先在 TiKV/Raft 事务语义中成立，然后 TiFlash 通过 Raft log apply、row-to-column decode、DeltaMerge 写入、segment snapshot、DeltaIndex 等步骤把数据转成可分析查询的列存视图。因此一个读结果不一致，可能发生在事务层、Raft learner apply、proxy 交互、行转列、DeltaMerge 存储视图、查询执行任意一层。相比单一存储引擎，怀疑面更宽。</li><li>TiFlash 的“读到某个 TSO”不是简单读一个本地 MVCC store。它需要 ReadIndex 保证 learner 至少 apply 到足够新的 Raft index，还要处理 Async Commit lock、<code>max_ts</code>/<code>min_commit_ts</code>、read through lock、bypass lock、resolve lock、Remote Read 重试。也就是说，读可见性由 TiKV 事务协议、Raft read index、TiFlash learner 状态和本地 DeltaMerge snapshot 共同决定。任何一层日志单独看都可能“合理”，但组合起来才暴露矛盾。</li><li>Region Split 和 Remote Read 会把一个逻辑 table scan 拆成跨 region、跨 TiFlash 节点、跨 epoch 的多个子读。Issue #8845 的现场多次出现 split 后 epoch mismatch，随后触发 Remote Read；Remote Read 又可能在另一个 TiFlash 节点上以 Coprocessor 请求形式读同一个 region。这样一来，同一个 SQL 的错误可能只出现在某个 region 的某个 remote cop task 上，而其他 TiFlash 节点和后续查询都正确。定位时必须把 SQL、MPP task、region epoch、remote target、read index、segment 读任务全部串起来。</li><li>TiFlash 本地列存不是 Raft KV 的直接镜像，而是 DeltaMerge 的多层读视图。一次读会拿 segment snapshot，再决定是否复用 shared <code>DeltaIndex</code>、是否需要 <code>ensurePlace</code>、是否 <code>fully_indexed</code>。这类内部优化通常不改变持久化数据，只影响某个瞬间的读视图，所以 bug 表现为“瞬时少读，后续恢复”。这种问题不会像数据损坏那样留下稳定现场，也不会像事务协议错误那样一定能从 TiKV 日志直接解释。</li><li>TiFlash 为了性能引入了多处跨请求复用：ReadIndex 结果复用、shared DeltaIndex 复用、batch read index、remote read 重试、snapshot 复用/推进。这些优化单独看都有正确性条件，但一致性 bug 往往出现在“两个正确优化的边界交错”上。#8845 就是旧 snapshot 和过新的 shared <code>DeltaIndex</code> 在 duplicated tuple 场景下组合出了错误读视图。</li></ol><p>和其他 HTAP 实现相比，这个复杂度有 TiFlash 架构自身的特点：</p><ul><li>SingleStore 官方文档强调它可以在一个系统中结合 rowstore 和 columnstore，并在单个查询里合并实时和历史数据：<a href="https://docs.singlestore.com/db/v9.0/introduction/how-singlestore-works/high-performance-for-oltp-and-olap-workloads/" target="_blank" rel="noopener">High-Performance for OLTP and OLAP Workloads</a>。它也有分布式和存储格式复杂性，但通常不是“TiKV 行存主副本 + TiFlash Raft learner 列存副本”这种跨进程、跨引擎复制后的读视图问题。</li><li>CockroachDB 更偏向单一分布式 KV/MVCC 存储层，上层有向量化/列式执行来加速分析型查询：<a href="https://www.cockroachlabs.com/docs/stable/architecture/overview" target="_blank" rel="noopener">Architecture Overview</a>。这类架构的难点更多集中在事务、租约、range、closed timestamp 等同一存储系统内的一致性，而不是额外列存 learner 副本与行存主副本之间的视图对齐。</li><li>OceanBase 官方 HTAP 介绍强调在同一个分布式数据库内直接承载交易和分析：<a href="https://oceanbase.github.io/docs/user_manual/operation_and_maintenance/en-US/scenario_best_practices/chapter_03_htap/introduction" target="_blank" rel="noopener">OceanBase HTAP introduction</a>。它同样有分布式事务和分析执行复杂性，但 TiFlash 这种“主行存 + 异步列存副本 + Raft learner + 行转列 DeltaMerge”的链路，给排查额外引入了副本同步点和列存视图构建点。</li></ul><p>所以 TiFlash 的一致性问题难查，并不是因为某个模块特别脆弱，而是因为它把事务系统、Raft 复制、异步列存副本、MPP/Remote Read、DeltaMerge 本地读优化组合在一起。问题的真实根因可能在任意一层，但表象往往会穿过多层后才被 SQL 对拍发现。#8845 正是这种架构复杂性的典型案例：外观看起来像事务/ReadIndex 问题，实际是列存读视图中 shared DeltaIndex 和 snapshot 的代际错配。</p><h2 id="关联-PR-和作用"><a href="#关联-PR-和作用" class="headerlink" title="关联 PR 和作用"></a>关联 PR 和作用</h2><ul><li><a href="https://github.com/pingcap/tiflash/pull/9000" target="_blank" rel="noopener">#9000</a>：最终修复。核心是阻止旧 snapshot 复用包含其未覆盖 duplicated tuples 的过新共享 <code>DeltaIndex</code>，并补充单测 <code>DupHandleVersionAndDeltaIndexAdvancedThanSnapshot</code>、<code>ReadWithMoreAdvacedDeltaIndex2</code> 等。</li><li><a href="https://github.com/pingcap/tiflash/pull/8873" target="_blank" rel="noopener">#8873</a>：排查期间改进 ReadIndexWorker 代码和日志，记录 read index 复用等信息。</li><li><a href="https://github.com/pingcap/tiflash/pull/8874" target="_blank" rel="noopener">#8874</a>：增加可选的 record version 日志和 segment snapshot rows 日志，用于确认写入版本、行数和 snapshot 关系。</li><li><a href="https://github.com/pingcap/tiflash/pull/8928" target="_blank" rel="noopener">#8928</a>：排查 remote read / async task 相关现象时用于检查测试路径，不是根因修复。</li><li><a href="https://github.com/pingcap/tiflash/pull/3971" target="_blank" rel="noopener">#3971</a>：历史上引入高并发下按 start-ts batch/reuse read-index 的机制，一度被怀疑会导致旧 TSO 复用过小 read index；最终由 <code>max_ts</code> 逻辑和后续证据排除为主因。</li></ul><h2 id="如果用-AI-辅助调查，能加速哪里"><a href="#如果用-AI-辅助调查，能加速哪里" class="headerlink" title="如果用 AI 辅助调查，能加速哪里"></a>如果用 AI 辅助调查，能加速哪里</h2><p>这类问题里，AI 最有价值的地方不是直接猜根因。早期现象确实很像 ReadIndex、Async Commit、lock 或 Remote Read 问题，AI 如果只根据表象下结论，也很可能走同样的弯路。真正能加速的是把大量日志、代码路径、假设和反证组织起来，让调查更快收敛到“哪个假设还站得住”。</p><ol><li>日志归并和时间线重建<br> 这个 bug 的证据分散在 endless 报错、TiFlash MPPTask、LearnerReadWorker、RemoteRequest、DeltaMergeStore、Segment、Proxy、TiKV raftstore 日志里。<br> AI 可以围绕同一个 <code>start_ts</code>、<code>region_id</code>、<code>applied_index</code>、<code>row_id/version</code> 自动抽取事件，重建“split -&gt; remote read -&gt; read index -&gt; resolve lock -&gt; raft apply -&gt; segment read”的时间线。</li><li>维护假设和排除矩阵<br> 排查过程中出现过 remote read key range、Region Split、apply snapshot、ReadIndex 复用、<code>max_ts</code> 未推进、memory lock、bypass lock、DeltaIndex 等多个假设。<br> AI 的作用是能够快速地去探明一个可能的原因是否是真的 root cause，从而去简化上下文。这是因为随着查询链条的深入，也引入了多个分叉，链条之间的因果就不是那么牢靠。从一个角度能解释的通，但是从另一个角度可能这个原因就解释不通。<br> 更重要的是，它可以提醒“<code>fully_indexed=false</code> + 后续 shared delta index 被推进”本身就是存储读视图风险点，避免在事务层钻牛角尖，从而更早转回 DeltaMerge。</li><li>生成更精准的 instrumentation<br> 排查中实际合入了 #8873 和 #8874 来增强日志。AI 可以基于当前假设建议更有针对性的打点，这能减少“先广泛加日志再人工筛”的迭代成本。</li><li>从现场日志反推出确定性单测<br> 最终 #9000 的测试本质上是把线上交错压缩成一个可控序列：线程 A 先拿旧 snapshot，线程 B 先读并推进 shared delta index，然后线程 A 再用旧 snapshot 读。AI 很适合根据日志里的 happens-before 关系生成这种 test skeleton，帮助把偶发线上问题变成稳定单测。</li><li>跨组件语义校验<br> <code>ReadIndexObserver</code>、<code>MsgReadIndex</code>、<code>RaftCmdRequest</code>、Async Commit、<code>max_ts</code>、<code>min_commit_ts</code>、read through lock、bypass lock 这些概念很容易混在一起。AI 可以把设计文档、代码和日志放在一起核对：当前路径到底是否推进了 <code>max_ts</code>，哪个 message 类型生效，<code>bypass_locks=[]</code> 能排除什么，<code>commit_ts &lt; read_ts</code> 的日志是否一定表示事务层错误。这个能力能帮助更快识别“看似矛盾但其实不是根因”的证据。<br> 这能够减少查问题的时候，因为涉及到不同组件，不同人之间的协调产生的损耗。</li></ol>]]></content>
    
    
    <summary type="html">&lt;p&gt;以几年前的一个 case 为例，讨论了 TiFlash 的 HTAP 架构对排查不一致问题的影响，也讨论了如何用 AI 来减少这一类问题的调查时间。&lt;/p&gt;</summary>
    
    
    
    
    <category term="数据库" scheme="http://www.calvinneo.com/tags/数据库/"/>
    
  </entry>
  
  <entry>
    <title>What&#39;s left for the engineers in the AI era</title>
    <link href="http://www.calvinneo.com/2026/03/24/new-engineer/"/>
    <id>http://www.calvinneo.com/2026/03/24/new-engineer/</id>
    <published>2026-03-24T15:20:37.000Z</published>
    <updated>2026-03-24T05:17:14.103Z</updated>
    
    <content type="html"><![CDATA[<blockquote><p>这是一篇简单直接的文章（或者被认为是提纲），目的是提出价值判断、提出关键问题。我刻意避免引用例子，或者论述逻辑来佐证观点，也刻意避免提出解决方案，因为我认为这一部分工作已经被 AI 所取代。</p></blockquote><p>本文主旨在于探讨在 AI 时代，一个工程师还能够输出什么样的价值。<br>《空中浩劫(ACI)》系列是理解工程思想很好的学习案例，而计算机系统和航空产业又有很多的共同点。下面几点尤其值得关注：<br>1.所有的黑天鹅都是灰犀牛<br>2.本质上都不相信人和机器，对环境保持最低限度的假设<br>3.计算机（AI、自动驾驶）的参与度都很高</p><a id="more"></a><h1 id="管控风险"><a href="#管控风险" class="headerlink" title="管控风险"></a>管控风险</h1><blockquote><p>瑞士奶酪模型（Swiss Cheese Model）：在复杂的工程系统中，灾难极少源于单一故障，而是多层防御体系中的隐患恰好‘对齐’所致。</p></blockquote><p>从单个模块的视角观察，一些缺陷的影响并不显著，但是在系统中则会成为致命风险中的一环。在实践中，这样的缺陷可能还会被放大，呈现“蝴蝶效应”。一个超出 SOP 的行为可能产生远超自己预期的结果。<br>这个模型没有强调的一点是，在计算机或者民航这样的高强度运行的系统中，一切有可能发生的问题最终一定会发生。因此，管控风险不能功利主义地乘上概率，而是真的要把那微小的 corner case 当做正常情况来处理。<br>我认为一个系统的复杂度和风险无论经过怎么样的封装或者 trade off，都是不能最终被消除的。既然风险是正常情况，管控和确认风险将始终是每个工程师的任务。</p><h1 id="和机器相处"><a href="#和机器相处" class="headerlink" title="和机器相处"></a>和机器相处</h1><h2 id="人机接管边界"><a href="#人机接管边界" class="headerlink" title="人机接管边界"></a>人机接管边界</h2><p>控制权在人与机器之间切换的那个“边界时刻”，往往最容易引发问题。此类问题的特征是机器或者人类都可以独立完成任务，但是当两者合作时就产生了交接过程中的掉棒，或者合作过程中的冲突。<br>这种问题本质上是在几万年的进化中，人和人之间已经形成了诸多共识；在机器设计的过程中，各个模块间指定了 protocol，并通过测试来保证正确性。但是人和机器的交互之间并没有形成相互的理解，也没有信任。</p><h2 id="从写认知到读认知"><a href="#从写认知到读认知" class="headerlink" title="从写认知到读认知"></a>从写认知到读认知</h2><p>读优化和写优化是数据库系统中常常要思考的问题，对于人脑也是如此。从我的观察来看，人脑是一个写优化的系统：布道远比学习要容易，输出自己永远比理解他人要容易。</p><h2 id="统计学与偏见"><a href="#统计学与偏见" class="headerlink" title="统计学与偏见"></a>统计学与偏见</h2><blockquote><p>人类的大脑是一台出色的“模式识别机”，但绝对不是一台合格的“统计计算机”。 —— 丹尼尔·卡尼曼</p></blockquote><p>在我的文章<a href="/2021/05/15/probability-problems/">《概率论中的几个有趣问题》</a> 中论述了一些“反直觉”的统计问题。<br>但更本质的是，无论是“厌恶”还是“偏见”，人类更愿意基于经验去得到一个非 0 即 1、非黑即白的结论，这是厌恶非确定的本能。关于这个问题我在 <a href="/2025/02/09/logical-fallacies/">《形式谬误和非形式谬误》</a> 一文中也进行了详细的论述。<br>我认为 AI 的出现会加剧这样的问题，因为 AI 迎合了人类这样的性格弱点。</p><h2 id="从提出问题到构建价值"><a href="#从提出问题到构建价值" class="headerlink" title="从提出问题到构建价值"></a>从提出问题到构建价值</h2><blockquote><p>Code is cheap, show me the talk.</p></blockquote><p>在 AI 年代，这句话指出了提出问题比解决问题更加困难也更加有价值这个事实。因此本文的写作中，根本也没有提出解决方案，毕竟有太多关于它们的讨论了。</p><p>但随着 AI 将能力边界扩展到了“主动式智能”领域，它们是有能力提出一个好问题，从而推动进化和进步的。所以问题留给了人类，当 AI 澎湃的生产力剥夺了人的利他性，那么人类的最终价值是什么？</p><p>我想这个答案应该很简单，当人的所有属性都被剥夺之后，人的价值就不再依赖任何的属性和定位，人的价值就在于他是人。我希望 AI 能够引导一个伟大的社会变革，但留给人类最应该做的是为 AI 的发展方向提供一个优化目标，也就是构建价值。</p><h1 id="比失败更糟糕的是“局部成功”"><a href="#比失败更糟糕的是“局部成功”" class="headerlink" title="比失败更糟糕的是“局部成功”"></a>比失败更糟糕的是“局部成功”</h1><p>最后想再多说一句，现在是一个人人 demo 的年代，”Done” 的门槛已经越来越低，Done is no longer better than perfect。<br>那么什么是真正的门槛呢？我想 BERT 本身就是一个很好的反例，它的失败恰恰可以归功于它的成功。也许我们每次在设计的时候，真的需要多问自己一句，这个架构是不是真的就止步于此了？</p>]]></content>
    
    
    <summary type="html">&lt;blockquote&gt;
&lt;p&gt;这是一篇简单直接的文章（或者被认为是提纲），目的是提出价值判断、提出关键问题。我刻意避免引用例子，或者论述逻辑来佐证观点，也刻意避免提出解决方案，因为我认为这一部分工作已经被 AI 所取代。&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;本文主旨在于探讨在 AI 时代，一个工程师还能够输出什么样的价值。&lt;br&gt;《空中浩劫(ACI)》系列是理解工程思想很好的学习案例，而计算机系统和航空产业又有很多的共同点。下面几点尤其值得关注：&lt;br&gt;1.所有的黑天鹅都是灰犀牛&lt;br&gt;2.本质上都不相信人和机器，对环境保持最低限度的假设&lt;br&gt;3.计算机（AI、自动驾驶）的参与度都很高&lt;/p&gt;</summary>
    
    
    
    
    <category term="Articles" scheme="http://www.calvinneo.com/tags/Articles/"/>
    
  </entry>
  
  <entry>
    <title>呼和浩特和大同游记</title>
    <link href="http://www.calvinneo.com/2026/02/24/meet-in-hohhot-datong/"/>
    <id>http://www.calvinneo.com/2026/02/24/meet-in-hohhot-datong/</id>
    <published>2026-02-23T17:20:33.000Z</published>
    <updated>2026-02-25T14:28:15.425Z</updated>
    
    <content type="html"><![CDATA[<p>过年期间去了呼和浩特和大同。这是首个依照 codex 做的旅游规划玩的活动。</p><a id="more"></a><h1 id="D0"><a href="#D0" class="headerlink" title="D0"></a>D0</h1><p>我的行程在 <a href="https://github.com/CalvinNeo/CalvinSchedule" target="_blank" rel="noopener">https://github.com/CalvinNeo/CalvinSchedule</a> 中。</p><h1 id="D1-初一"><a href="#D1-初一" class="headerlink" title="D1 初一"></a>D1 初一</h1><p>因为我对象年三十值班，所以我们是定了初一下午的飞机飞呼和浩特。为什么去呼和浩特是因为去长春的飞机票买晚了，所以在选择从大连出发玩东北，还是从呼和浩特出发玩呼和浩特和大同之间的权衡。</p><p>事实证明，呼和浩特是一个挺好的地方，首先它去大同比太原还要方便，其次，它的机场到高铁站很方便，所以我们直接定一个高铁站旁边的酒店就非常舒服了。不过据说呼和浩特的新机场要启用了，因此后面它就没太原爽了。</p><p>因为我们值机比较靠前，所以下飞机是比较快的，然后就直往地铁站跑，赶上了倒数第二班地铁。事实证明坐地铁还是对的，因为后面我们点了个鼎鼎眼镜烧烤，发现等了好久都没有骑手接单，最后还是加钱人家才送的。</p><h1 id="D2-初二"><a href="#D2-初二" class="headerlink" title="D2 初二"></a>D2 初二</h1><p>今天起来还是很感冒，还是流鼻涕。</p><p>早上 8 点，我们的小团才接我。这也是因为我们住在呼市最东边的缘故，所以不用起那么早。车上除了我全是女的，所以我就坐在副驾驶了。这个副驾驶的头枕非常靠后，所以我睡起来很不方便。</p><p>第一站是辉腾草原上面的马场。这个马场除了骑马和卡丁车（实际上没人玩）啥都没有，所以尽管我们没有定骑马的套餐，所以逼到最后还是只能骑马。那边报价 140，而淘宝只要加 120，所以我们用淘宝价还价骑了马，然后又加了 120 骑了快马。<br>大概流程就是先上一匹小马，然后几匹马一起被牵着走到跑马场。到那边老板就问要不要骑快马，如果骑，就换到大马身上。然后老板就会同时拉着你的缰绳和自己的缰绳把马跑起来，嘴上还嘚嘚嘚嘚的。一开始马是小跑，这个时候你得用腿用力夹住马的肚子。不过后面马就会撒开来快跑了，这个时候，就是颠屁股了，感觉整个人都要散架，必须要拉着马鞍才会稳一点，然后就跑完了。</p><p>总的来讲感觉价格偏贵。体验上的话，快马还是可以的，不过时间也很短。当然，再长一点，我呛风也受不了了。然后这地方的旱厕也是堪称一绝，里面的屎冻成了棍子，我也是第一次见。导游说你就路边上尿吧，我寻思这么大的风，不吹到身上么。</p><p>然后就是去吃饭，开车到了右翼中旗县城里面，就开了那么几家店。我们选了一家，我对象点了个羊杂，点了个驼饼。我吃了下，感觉羊杂还挺香的，驼饼感觉就是普通的肉饼吧，然后查了下，说是骆驼肉。</p><p>吃完饭，就去火山了，这个也是要开一大段路的。那一块其实有很多火山，有的火山形状像草帽山，我把它叫土帽山。不过我们实际去逛的那个叫南炼丹炉，是一个平顶的火山。相比我们在冰岛看到的，这个火山确实有点火山的感觉。</p><p>从火山回来，就是长达三个小时的回程了。我们让司机帮我们送到宽巷子，这样就可以直接吃东西了。下车就能看到杨老大焙子店，结果没开门。旁边的星月排满了队，所以我们就去清和园了。这边买芝士奶饼好多都是几十个买的，我们只买了两个感觉很亏。然后我们就去珠萨拉买了几个酸奶糕，是冰着的，感觉挺好吃。</p><p>吃完就准备打车去泽成冰煮羊了，这个店在附近有两家。我决定去万象城那一家，因为那边等的时候至少可以逛一逛。事实上感觉也是对的，因为那天晚上我感冒贼难受，加上又很干，嗓子不舒服。排队也是一百多号，特别渴，所以想买个奶或者沙棘汁喝喝，找到一个牧野晨曦奶茶店，说是要等一个多小时，立马吓退。</p><p>反正最后总算是进去了，这个冰煮羊必须要至少点一斤的羊肉或者羊排。</p><h1 id="D3-初三"><a href="#D3-初三" class="headerlink" title="D3 初三"></a>D3 初三</h1><p>今天需要早起去大同。因为我们在火车站旁边住，所以提前一个小时起床绰绰有余了。内蒙古的火车会说什么内蒙古好羊肉，内蒙古羊肉好，非常洗脑。反正一路上也是比较困，睡睡醒醒的。<br>从一路上看，内蒙古和山西确实不太一样，首先，内蒙古真的有很多风车。然后，内蒙古感觉是草原的感觉，是有黄草的，但山西的话会更有黄土高原的感觉。</p><p>下了大同站，想要拉屎。我对象不在车上拉，结果下了车就要排队。其实始发站的火车，我一般是喜欢在火车上拉屎的，因为一般都很干净。拉完屎还有五分钟，就跑到直通车那里，准备去云冈石窟。上了车，开个大概有一个多小时才到，真的是非常非常堵了。</p><p>云冈石窟人非常多，主要体现在：首先，它景区检票口开了好几个，但居然都要排队。其次，是它的讲解也要排队，我们想想还是旅途随身听算了。反正进去之后，就绕了一圈，看了一下它的一个寺庙的前中后殿，然后过了桥，就是石窟的主体。</p><p>因为我之前在车上已经预习过了，所以我大概知道这些石窟大概分为几个系列，然后哪些是比较精华的。反正一开始进去的第一窟和第二窟都是比较惨烈的，因为它们的开口很小，所以基本人只能站在石阶上往里面伸着头看，我似乎就看到几个石柱子，感觉没啥意思。</p><p>第五窟和第六窟是一起的，实际上我们排最长的也是第六窟，因为第五窟当时已经维修了。保安说要排一个多小时，但实际排了四十几分钟就进去了。这个窟实际上是有前窟和后窟的，云冈石窟大部分也都是这个式样。并且相比于莫高窟或者龙门石窟，云冈石窟的颜色没怎么掉，所以看起来更鲜艳，据说这个可能是因为之前有人补过颜色吧。</p><p>在排第 6 窟之前，还看到旁边有一个可以往上爬的地方，当时也没有注意是啥。回来的时候，我们找从上面下来的人问了下，说也是一些小石窟，所以就没有再上去了。后面到那个院史馆看了下，说上面好像是有个什么寺的遗址，我觉得也是挺神奇的。</p><p>9/10 两个窟或者是 7/8 两个窟好像是一对的，就是它们的构造都很相似。据说是什么太后要和皇帝平起平坐，所以搞了两个相同的。</p><p>11/12/13 这几个窟是连在一起的，也就是有名的音乐窟。这几个窟人不是很多，感觉是因为窟比较大，所以能够流水线进出。</p><p>14 窟看不出来是啥。15 窟是那个万佛洞，里面雕的全是佛像，看起来有点密恐。</p><p>16-20 就是著名的昙曜五窟了。</p><p>看完 20 窟，再往前走就是西部窟群，里面的东西就比较挫了。</p><p>旅途随身听上看到有个云冈石窟博物馆，加上之前听到说有什么东西被抠出来放到了博物馆里面，我就想去看看，结果这博物馆就没开。我记得之前好像也是哪个景区里面的博物馆也不开的，感觉还是比较遗憾的。</p><p>出去的路也是比较奇葩的，因为这个景区里面有一堆美术馆，但是都关门了，最离谱的是还有一个侏罗纪断层的展示，上面放了一堆恐龙模型，也不知道是在干嘛。</p><p>琥珀知道我们到了大同，就想跟我们面面基，一开始提议是去古城看灯展。所以我们从云冈石窟出来，就买了直通车的票，准备回古城，和琥珀约一下。这个直通车体验是真的差，首先，车里面暖气开的热的像桑拿一样。然后，路上又贼堵，关键是这车非要沿着堵麻了的平城街开到和阳门那里下客。我能理解上客得在那，但是既然那么堵，为什么不能在古城北门先让我们下一下呢？本来我们是准备坐这个车先到宾馆放一下东西再进古城的，但是这一下子开了快两个小时，感觉再回去就来不及排队吃饭了，所以我们就在古城下车了。一进古城就发现东门排着队，我们觉着应该是上城墙准备去看花灯的，所以也果断不去看灯展了。</p><p>琥珀显然对大同的人流没有预期，搞了半天说要吃紫泥，结果我们去古城紫泥拿了个号，发现要等 290+ 桌。不过我们倒是在古城的凯鸽也排了个号，只有 29。不过考虑到它一直不叫号，并且古城也没啥可以逛的（人特别多），所以我们就准备出去找点东西吃。经过一系列的电话咨询，我们发现紫泥在南环路的一家店还是可以取号的，所以就骑自行车过去。</p><p>骑车过去的路也是比较奇葩，主要是出了古城城门需要往右拐，结果交警让我们走上面走，结果骑着骑着就骑到了永泰门广场上了。这个广场是被拦着的，自行车没法进，我们也下不去。结果我们只能闯过一片烂泥地强行进去，然后又在前面的栅栏的地方，把自行车搬了出去。骑到紫泥那里，它家在三楼，但是只有直梯上去，感觉挺麻烦的。上去之后果然全是人，反正是取了个号，然后打算先去找点东西垫垫肚子。先骑车去旁边的和笙财，点了个沙棘冰淇淋，这个冰淇淋是真好吃。然后发现旁边就是老柴削面的总店，但是我对象还是想吃喜晋道，所以就骑车去旁边的喜晋道，中途路过大同一中，感觉这个学校真的是富丽堂皇啊。进去喜晋道，他家是一个非常大的门面，装修非常漂亮，但是一进去，就说没位置了，并且都不发号了。于是就灰溜溜排老柴削面。再过去，就发现老柴削面外面都全是人了。因为反正没事做，就进去点了个最基础的猪肉面，后面就是漫长的等待。我对象要找个座位，我是受不了，因为里面又热又闷又吵又挤，感觉就像是浴室里面一样。然后不断还有人端着碗要你让，或者叫着“臭油”把一堆泔水一样的东西拿到外面倒掉。我一直在外面等，发现有人把盘子端到旁边的黄焖鸡的店面里面吃，黄焖鸡的老板也是日了狗了。这家店外面还有叫什么老柴门口烧烤，也有人点，不过我搜了下，说这烧烤贼一般，就算了。</p><p>等老柴削面的过程也是很累，我们应该是 834 号，点完单的时候刚叫到 78x 号。然后一堆人就在那个打面的地方排着听叫号。这里面坑的是一个号可能点很多碗面，比如一个东北大哥一个号点了七八碗面还有一堆小菜，前后端了大概有五六次。相比之下，我们这一碗面就显得很亏。反正最后还是等到了面，吃起来，感觉挺一般的，有点类似于之前在太原全季早餐吃的刀削面。也许有那么一点更劲道，或者汤底更香吧，但可能是纯感觉，唯一可以确定的是这汤底是确实齁咸。</p><p>吃完了老柴削面，紫泥差不多也快到了，于是就走路过去。上去之后，发现大厅等得全是人，然后还有人源源不断上来，然后服务员说还要再等两个多小时。我们好像是 179 号，反正就是硬等了。有一个大哥盯着监控，看哪一桌走了，就说我就是下一个号，你给我就安排那里吧，感觉是饿坏了。很快到我们了，然后我们前面的号问我们要不要，自己当时排了两个号，我说我就是下一桌了。</p><h1 id="D4-初四"><a href="#D4-初四" class="headerlink" title="D4 初四"></a>D4 初四</h1><p>今天继续是早起的一天，我们需要去恒山景区。从酒店到南站有一定距离，还是扫了个车骑了过去。然后，发现百度导航真的是糟糕，直通车明明是车站的另一端，它给标错了。总之赶到了直通车那，说恒山的车堵在高架上，可能要晚点。然后我就去上了个厕所，果然是上厕所定律了，上到一半，车就来了，结果我赶快跑了回去。在车上迷迷糊糊睡到恒山，到了那个游客中心附近又开始堵车了。总之到了下面，就要坐摆渡车，我也看到了之前在小红书上看到的排队标志，所幸今天的人没有昨天多，不过悬空寺方向还是挺长的，所以我们决定先去爬恒山了。</p><p>去恒山的车会经过悬空寺，在经过一个隧道之前，能看到悬空寺，并且这是一个很好的从上往下看的机位，建议不要错过，因为返程的时候，需要倒着头向后来看，不是很方便。</p><p>到了恒山脚下，我决定还是坐摆渡车上山，爬上去，然后索道下来。中间我对象买了个帽子，然后就又去排上山的队。今天可能景区被骂优化过了，所以我们在排摆渡车之前就检票了，没有等很久。一辆车走了之后，等了一会，然后突然开过来五辆车，都并排停了。我们上了第一辆车的最后一排，我对象不太乐意，因为她有点晕车。</p><p>上了摆渡车，很快就到真武庙了，从这里就往上爬。我带了个 insta360 的相机，可以看到，全程都很简单，我觉得都没有南京紫金山难爬。中间有个庙，爬上去的石阶比较陡，感觉应该是最困难的部分了，但其实也就那样。然后从庙出来，就是冲顶阶段了，过了一个叫氵麦极门的地方之后，就看到小红书上堵人的地方，我们不出意料也开始堵了。感觉就是从这里开始，台阶变得高了点，所以有的人就要歇歇了，加上石阶又变窄了，所以交通就阻塞了，好在这一段不是很长，很快就到了一个亭子那。那是一个三岔路，往左是索道下山方向，往右是登顶方向。我们登顶，好在那里石阶就宽很多了，基本都可以跑起来，所以跑了大概十几分钟，我也登顶了。总共算上等的时间是六十几分钟吧。总的来说，恒山其实没啥意思，主打一个五岳打卡。然后确实很多人都在峰顶的那个地理标识那边排队拍照，说那边一个人只有 20s 的拍照时间。</p><p>下到缆车的地方，看到人也不是很多，就打算坐缆车，结果下来之后发现别有洞天，还是排了不少的。不过其实是可以忍受的，实际上也就等了不到二十分钟。中间前面的京爷因为不知道什么事情吵起来了，然后又打起来了，然后又吵起来了，反正我们就翻栏杆排到了他们前面，感觉也是个乐子。下山就直奔摆渡车站，去古城的在排队，但是去悬空寺的不需要排队，我们就直接上了。</p><p>很快就到了悬空寺，看到很多人在往里面走，好像是自驾过来的，所以我其实有点想看完悬空寺就直接打车回去浑源古城看永安禅寺，省得坐他那个摆渡车兜圈子了。anyway，我们往里面走，是茫茫多的人，然后进去也要走一段路。反正到了悬空寺下面，感觉其实进不进去看也一个样子了，不过来都来了，就进去看一下。里面就是一个小广场，给你拍拍照，然后有一个地方是登临的入口，感觉排了一堆人，我们估算了一下，感觉肯定不止景区宣传的 3260 人每天的量。然后这波人是一波一波往里面放的，所以大家都只能慢慢等，我挺庆幸自己没找黄牛买票，感觉过来排队也还是受罪。过了一个桥，可以更近看悬空寺，不过其实没啥区别，主要这寺确实很小。我对象说梁思成其实也不怎么 care 这个寺，我也不知道为啥。</p><p>出了门，问了保安，说这边社会车辆进不来，他也不知道哪里能打车。我打开滴滴，发现自己站的地方就能打车，所以滴滴肯定是在扯淡。既然打不了车直接去浑源古城，我们只能坐大巴先去游客中心了。快到游客中心我们就远远看到了电瓶车，等下了车，我们就扫了个更近的电瓶车，往浑源古城那边骑。这电瓶车还挺贵的，骑了几公里，扣了我四块钱。</p><p>到了浑源古城，先去永安禅寺逛了下，这个也是辽金时候的建筑，也是出现在浪浪山这个电影里面的。主要看的是它主殿的壁画，正面是画的各个法王，最有名的就是不动明王，他是一个白净的，撕开脸的造型，和其他的法王区别很大，所以他在文创上的露脸也很多。另外一个法王的眼睛很奇特，从左边看是紫色的，从右边看事棕色的，说是因为它眼睛是用贝壳做的，所以有不同的反光。然后左右两边的说是明清的时候画的，其中右边的好像是道教的一些人物，不太看得懂，左边的听了一圈解说，说是佛教里面的六道轮回。我们蹭了好几个解说，所以大概都听了个明白，其中有个解说甚至把谁是东海龙王，谁是地藏王都指出来了。讲解还贴心的照了一下壁画，给我们看这个壁画实际上是有厚度的。</p><p>从永安禅寺出来，就去吃了下小媳妇凉粉。浑源这边遍地都是小媳妇凉粉，但是我们吃的好像是真的小媳妇。点了个凉粉，还有个熏蛋。凉粉啥感觉记不得了，熏蛋还真的挺有意思的。喝了他们的沙棘汁，感觉不如昨天的紫泥的酸。</p><p>骑电瓶车回去，差不多要赶六点的直通车。发现根本找不到直通车的乘车点，百度地图真的是拉胯，但是切换高德地图也找不到。反正问了一圈，终于找到了一个标志，跟着走下去。发现除了 901，还有 25 块钱和 39 块钱不一样的直通车。因为我们已经买了 39 块钱的，所以就过去检票。然后因为我们买的是 7 点的，所以还涉及到改签的问题。第一班车已经只剩下最后两个坐了，所以我们跟检票的小哥一顿交涉，跟着前面几个妹子一起去坐下一班车了。然后小哥上来先让我买一个新的票，然后他再给我把老票直接退了。然后就是等，车很快又要坐满了，中间来了两个女的，发现又只有最后一排的坐了，就想等下一班，但是司机说又已经检过票了。反正折腾了一圈，她们去打车了。然后上来了几个，又有不愿意坐最后一排的，总归折腾了很久还是发车了。</p><p>我们选择在大同南站下车，因为我们路上排了两家凯鸽都要等一段时间，我希望去宾馆洗一下屁股，顺便看看有没有其他能吃的。这个时候琥珀也要来找我，说怎么说再吃个饭按摩一下。因为太饿了，我还请宾馆帮我热了下昨天的肘子垫吧垫吧，不得不说，那肘子还是可以的。然后我就让他来宾馆，自己也顺便洗了个衣服。然后他到了，我们发现凯鸽水泉湾店快排到了，然后就赶快打车过去。到了之后，我对象进去等，我们出去逛了下，发现旁边就是他们的御河，我理解就类似于太原的汾河吧，不过没有自行车道，感觉不如太原做得好。</p><p>凯鸽的菜：</p><ul><li>风沙鸡：感觉很一般，就跟我们这边普通的脆皮鸡差不多</li><li>糖醋里脊：很香，我理解主要还是山西那边醋确实强调香吧，然后就是也没有那么甜</li><li>炸油糕：我朋友说不好吃，不过我们点了个尝一下还不错</li><li>烧麦：皮跟纸一样，但吃起来感觉没啥意思</li><li>羊肉火锅：感觉一般，我对象说味道太重了，我倒是挺喜欢这味道，但是我觉得太肥了，有点我点的锅圈食汇的肥羊的感觉，不是很喜欢</li><li>沙棘汁：还不错，没有紫泥的酸，但是挺普适的</li><li>酸奶：挺好</li><li>烤骨头：有点里脊肉的味道，但是是肘子的感觉，挺有意思的</li></ul><p>一边吃，一遍搜附近的按摩，感觉这边按摩都挺喜欢擦边的。我觉得心思放在擦边上的店基本上手艺挺一般的，所以我直接过滤掉了。反正最后选了一家 150 的，其实做下来感觉也挺一般的。这家店开在小区里面，我们过去的时候，也是穿过小区走过去的。中间我们又看到了山西的省灯，我觉得他们是不是会暗自较劲，看谁家的灯最花里胡哨？反正到了那家店，感觉是个坑姐妹钱的店，全是美容的设备。我们的 80 分钟果不其然是只按 60 分钟，然后做个热敷 20 分钟。然后他们急着下班，又边做边热敷了。我做的是什么泥膜，感觉非常烫。我感觉自己可能都要被烫伤了，就让他给我摘下来，一摸发现全是黏糊糊的汗。因为一直在烫伤的那种疼，所以让她拍照片发现皮肤是白的，看来确实是没有烫伤。不过这个后劲确实大，我大概缓了有一个多小时，这个烫伤的感觉才消失。所以这样看来，感觉这个泥膜还是挺舒服的。</p><h1 id="D5-初五"><a href="#D5-初五" class="headerlink" title="D5 初五"></a>D5 初五</h1><p>今天算是比较奇妙的一天。首先，昨天晚上睡觉前跟我对象讨论了，说今天早上去应县木塔吧。然后晚一点我就想把直通车的票买了，结果一看，9.30 的票只有一张了。因为木塔下午人多，并且，从木塔回古城更顺，所以我们又只能买早上的，就很尴尬。我还不信邪，又刷了几遍，中间还错买了去恒山的，总归坐直通车是不行了。灵机一动，决定去搜下有没有其他的巴士，结果 902 好像要两个多小时，不过中途发现，应县是有高铁站的，并且高铁站也是有巴士去木塔景区的，所以我就决定买火车票了。结果 12306 晚上关门，没法买票，于是我只能定了早上的闹钟。早上起来，发现买票是需要在系统里面排队的，我等了几十秒发现还没刷出来，就又小眯了会。几分钟又醒了，发现刷了两个不挨着的座位，果断付了款，结果再一开，火车票就卖空了。</p><p>结果上了高铁，发现这趟高铁也不算特别挤啊，不知道在那里票都卖光了。不过我们一站就到了应县西站，然后门口就是蓝色巴士，五块钱就到木塔了。然后发现这个巴士似乎也接路上的村民，并且送到木塔之后，也会继续往前开的。</p><p>因为我们是高铁来的，所以应该比大部队要早，到的时候，木塔没啥人，跟之前小红书上看到的完全不一样。我们进去之后，逛了两次一层，分别是从左边和右边走的，所以对里面的雕塑和浮雕看的都比较清楚。印象里面，可以看到木塔是有梯子可以上去的，不过特别陡。然后佛像正面是有两个洞，说是之前被破坏过。</p><p>实际上我们去的时候，木塔二层上是有工作人员在跑动的，只是不知道他们在做什么。站在木塔下面，如果角度不对，其实是不太看的出来木塔已经歪了的。但是确实能看到它的暗层，也确实能看到之前被敲掉的泥墙被换成的窗户，也理解那边的本地人为什么要把泥墙砸掉，因为换成窗户确实好看。</p><p>木塔一个拍照片的角度，就是他的几个底座的角落，我们在那边拍了照片，感觉挺好。不过如果要看木塔，我觉得还是站在下面比较好，我觉得是非常漂亮的，有种雍容华贵的感觉。特别是当你看到木塔周围那些拆迁到一半的破落房子和杂物堆的时候，这个木塔就更有一种清水出芙蓉的感觉了。</p><p>准备离开木塔的时候，风突然大了起来，并且沙子也多了。身边的游客在感叹，这地方还真的是太干旱了，风沙多，结果后来才知道，这是多年没见过的沙尘暴，是从蒙古国吹过来的。总而言之，离开的时候，回望木塔已经是灰黄笼罩的一片了。</p><p>前面那条街还在表演，看了会，就往回走，结果那妖风是越来越大了。走到南门出来的哪个巷子中，推开一个铁门，然后就走到一片很很乡下的地方。前面几个人躲着不肯走，我们过去一看，好家伙，前面黄沙都飞起来了。因为我们赶时间，就硬走过去了，我对象新买的帽子不知道什么时候被刮飞掉了。我们来回找了一下直通车在那里，因为之前在刮风那里确实远远看到“景区”几个字，但是走过去发现不是景区直通车，而是景区官方什么什么的。反正最后问了人，确定了就是在净土寺的广场，和我在微信公众号上看的一样的，所以我就直接导航到净土寺，总算是赶上了。不过上去之后，又是最后一排的位置。不过幸好之后还有一堆人要上车，所以最后我们换了个更大的车，因此又做到前面了。</p><p>上车发现琥珀发的消息，说是已经登机了，结果是什么沙尘暴，说又被赶下飞机了，气的要死。难怪我们这全是沙子，我先前以为他们这每天都这样呢，心想这样也太不好过了吧。</p><p>下午的计划是去古城看华严寺和善觉寺。我们先到古城看看晋喜道的情况，当然是停止放号了，跟我们说晚上的号下午四点开始发，要我们到时候来取。然后就网上搜了下刘姥姥豆面，发现在那个全是牌坊的十字路口，于是就过去。因为中途路过北魏家宴，所以我就说我们也去取个号，我对象说包取不到的，我就不信邪，结果还真只要排一小会，所以我们就先去吃了刘姥姥豆面，顺便等北魏叫号。</p><p>BTW，从下车的那一刻起，我们就感觉到了这个沙尘暴的厉害。进入古城，发现里面人少了很多，估计是风太大都不敢出来了。古城里面的灯会据说也是取消了。走到那个牌坊路口的时候，感觉自己仿佛到了黑沙滩，这风打在脸上是真的疼。琥珀说他的飞机已经延误到三点多了，然后我说你能不能航旅纵横买一个延误险，我看这风，你们三点估计都起飞不了，他说买不了了。实际上后面看小红书，就是我们快到古城，然后下车，也就是 12 点到 1 点那阵子风沙是最大的。</p><p>回到刘姥姥面馆，这面馆也是一个要看本事 DIY 座位的地方。不过好处是它有夹层和二楼，所以我们在二楼找到了个位子。我们点了一些小菜，以及一碗面，里面有它的豆腐干，还有两条青椒。</p><p>吃完刘姥姥，看到北魏的号是刷刷刷地过，突然就只有三个了。我们赶快跑过去，我对象还摔了一下。到了的时候，只有一个号了，所以我们就可以预点菜了。</p><p>吃完北魏，就去华严寺了，这又是蹭解说的一天。一进去，感觉里面鲜花盛开，我以为是冬天快走了，开了一堆梅花，结果走近一看是假的。华严寺一个看点就是大雄宝殿，这个大雄宝殿有一对超级大的鸱吻，另一个特别的是它的入口的门造型比较奇怪，据说是借鉴了伊斯兰教的一些风格。这个大雄宝殿里面到底是啥，因为里面很挤，解说声音又很小，所以我也没啥印象了。然后是普贤阁和文殊阁，好像一个是新的，一个是旧的。我看牌匾全是颜真卿题的，但看这个样子，感觉不是什么老物件。看完大雄宝殿就准备走了，但是发现那边还有个华严塔，然后那个塔排了好长的队。一查，原来是有个纯铜打造的地宫，于是决定还是排吧。当时风还是挺大的，所以我们顶着风排了好久，我中途还要上厕所也是忍下来了。这玩意要排队是因为它要换鞋套，这就慢了。然后首先是下地宫，这个地宫还是不小的，里面有一个台子，供奉了舍利子。相比南京的大报恩寺，这个舍利子还是挺慷慨的，能够够着看到，感觉就是个比较长的米粒一样。逛完了地宫，就往上走，爬塔。据说这个塔是仅次于应县木塔的塔，因为应县木塔不让爬了，所以要爬这个塔。爬的感受就是小号的姬路城，我觉得比较好的是爬的过程中可以看到塔的暗层。暗层实际上有很多像倒金字塔一样的结构，不知道是干嘛的。塔总共有三个明层，上面没啥好看的，风景也一般。下来之后，就准备出去了，半路上看到一个薄伽教藏殿，这个殿是个比较屌的殿。里面全是佛像，佛像上全是灰，解说的意思是这个灰从佛像造完之后，就没有被清理过了。里面有一个佛像是一个露牙齿的女的，说是照着一个小女孩做的，反正我也不是很懂，反正梁思成说很屌，有个什么称号。</p><p>逛完华严寺，刚好四点多了，赶快出去抢号。结果发现晋喜道拿号居然也要排队，我们拿的是 c90，取号的小哥说也不知道要排到啥时候，反正五点钟才开始叫号。所以我们赶快去逛一下善化寺，因为善化寺在南门，所以我们要走好远。基本上就是第一天我们到大同古城走的路线了。反正到了之后，善化寺门口有一个五龙壁，说是之前什么寺拆掉了，然后移动到善化寺门口的。进去之后，当然又是蹭导游了。善化寺的一个重点是它的三圣殿，这个殿比较有名的是那个像菊花一样的斗拱，说这个斗拱即使是辽金的建筑都很少见。然后它的大雄宝殿里面也是一些塑像，有一个塑像比较有意思，是说一个小鬼被点化成为一个善良的送子娘娘，然后就把他们一起塑像了，说是很少见的。然后讲解还介绍了一下工艺，说是塑像是整体镀金的，然后需要显示金线的地方，就把颜料扣掉。另一个塑像栩栩如生，讲解说他腿部的青筋都被雕刻出来了。这个殿里面还有个碑，说是朱熹的什么亲戚的，有一个讲解还指着说第几行有什么字，然后说了什么事情，我是没搞懂。总之觉得这两个殿，旅途随身听的讲解做的比较烂，得一直蹭讲解才行。</p><p>在排队的时候，我要上厕所，结果用百度找了几个厕所，一个都没有用，还是问的别人。后面立即下了高德，发现高德是准的。上完厕所，就直接去晋喜道了，当时已经快排到我们了。我们点了个肉末的面，一个番茄面，几个小菜，鸡爪，串和沙棘汁。小菜方面，我觉得山西的醋确实香，配上大蒜末，所以那个蒜泥肘花就很好吃了。鸡爪感觉就是茶叶蛋的味道。凉拌黄花算是山西特色了，和其他店里没啥区别。肉末面和昨天吃的老柴总店也是比较类似的，不过没有那么咸，更有点汤面的感觉。番茄面比之前在全季早饭吃的稍微好点，主要它不太是那种面是面，浇头是浇头的感觉，相对来说入味点，番茄泥也弄得比较干净。</p><p>因为这两天走的屁股都肿了，所以吃完走到东门打了个车，就直接回去了。中途看到了九龙壁，隔着栅栏看挺漂亮的，结果进去还得买十块钱的票，那谁还进去看啊。等车的时候发现和阳门对面的城楼，也有一种晋祠的感觉。</p><h1 id="D6-初六"><a href="#D6-初六" class="headerlink" title="D6 初六"></a>D6 初六</h1><p>今天在大同逛一下博物馆，就可以回呼和浩特了。早上起来退房，打车去博物馆。我们没有抢到预约，所以只能买特展的票。一开始我以为可以在门口给大爷看完就退了，结果发现这只是一个预检票，后面还有一道需要刷二维码的闸机呢。</p><p>进去之后，逛了一会一楼，上了个厕所，出来然后一个女的就问我们要不要拼团讲解，说是 100 块钱，看我们可能要急着走，就说 80 块钱算了，结果我们就定了。中间想先去看一半穆夏的特展，结果人家说只能进去一次，所以就放弃了，又下来。不过好在那讲解员拉客能力比较强，很快就齐活了，她先照着大厅里面的那个壁画讲了一遍，后面我们才知道，这个壁画并不被公开展示。然后我们就又从恐龙那边开始逛了。不过不同的讲解员的路线略有区别，所以我们这次就没有讲那个编织壶。我问了下，那个恐龙为什么那么完整，讲解员说这个叫什么天镇恐龙，当时挖出来也不是完整的，说是考古学家拼的。反正没懂我意思，因为有些恐龙是会缺几个骨头的，但是这几个恐龙是完整的。</p><p>二楼是比较精华的部分，即北魏展厅，因为大同直接是北魏的都城。其实它有两个展厅，讲解员只带我们逛了第一个。第一个中又分为两个主题，第一个主题是一个墓葬里面的发掘物，包含了漆画屏风以及它的附属，包括挂它的架子、柱子、石墩子都展示出来了。然后还有对应的墓志铭，以及一些陪葬的陶俑。第二个主题是丝绸之路，包含一些陶俑、壁画以及玻璃制品。这里的玻璃制品还是很漂亮的，是非常剔透的蓝色。</p><p>三楼是辽金和明清展厅，大同是北魏的陪都。然后，看到了一个非常大的鸱吻，说是从某个大殿上拿下来的。说鸱吻是龙和鲸鱼的后代，它长得就很像鲸鱼的尾巴。</p><p>讲解团解散之后，就去逛了下穆夏的特展。这个人是画板画的，不过眼睛画的很传神。另外，他也会画油画。看完特展，我发现二楼其实还有个展厅没看，所以就又下去看那个展厅。</p><h1 id="D7-初七"><a href="#D7-初七" class="headerlink" title="D7 初七"></a>D7 初七</h1><p>今天早上打车去了大召寺。这个寺是藏传佛教的，我不太懂，感觉没啥意思。有一个乃琼庙，里面都是骷髅头，很吓人。</p><p>从大召寺出来，就是塞上老街，我们又去吃了泽成冰煮羊，吃完，就打车去内蒙古博物馆。这个车是真离谱，调个头就花了十几分钟，堵得要死，结果我们等了有 20min 才上车。</p><p>内蒙古博物馆是真的大，感觉建得跟人民大会堂一样，里面也很大。我们预约的早上的，结果下午去就刷不进去了，被门口的说了下，就放进去了。一楼的展厅没有时间看，就奔着二楼去了。二楼是一堆乱七八糟的展厅，影响深刻的有一个马年的特展，还有一个神舟飞船的展厅。另外，他有一个特别大的古生物的展厅，里面放了从寒武纪奥陶纪，到恐龙时代的各种化石。我怀疑这个博物馆跟云南哪里有什么合作，因为它展出了很多云南出土的文物。</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;过年期间去了呼和浩特和大同。这是首个依照 codex 做的旅游规划玩的活动。&lt;/p&gt;</summary>
    
    
    
    
    <category term="游记" scheme="http://www.calvinneo.com/tags/游记/"/>
    
  </entry>
  
  <entry>
    <title>Vibe 一个桌游模拟器</title>
    <link href="http://www.calvinneo.com/2026/01/30/vibe-open-board-game/"/>
    <id>http://www.calvinneo.com/2026/01/30/vibe-open-board-game/</id>
    <published>2026-01-30T15:09:06.000Z</published>
    <updated>2026-05-26T08:37:35.638Z</updated>
    
    <content type="html"><![CDATA[<p>作为一个桌游爱好者，我打算用 Codex 去 Vibe 一个桌游模拟器，这样我可以尝试自定义规则和 Bot 强度。这篇文章我会介绍我 Vibe 的经验。</p><p>我的项目是 <a href="https://github.com/CalvinNeo/OpenBoardGame%E3%80%82" target="_blank" rel="noopener">https://github.com/CalvinNeo/OpenBoardGame。</a></p><a id="more"></a><h1 id="D0"><a href="#D0" class="headerlink" title="D0"></a>D0</h1><p>作为 Demo 实现了一个掼蛋游戏。 </p><p>发现问题：</p><ul><li>AI 对规则理解非常不正确，例如缺少对三带二、同花顺炸弹的支持，并且也不能顺子的长度和一手牌数量的上限。</li><li>AI 对空间感不熟悉，一些提示文字和牌重合。</li></ul><h1 id="Jan28-Jan29"><a href="#Jan28-Jan29" class="headerlink" title="# Jan28 - Jan29"></a># Jan28 - Jan29</h1><p>在这 2 天中，我大概用了 45 刀左右的额度。完成了 Cabo、骷髅牌、你画我猜、璀璨宝石四个游戏逻辑和 Bot 的开发。并且，我还支持了断开重连、房间管理等机制。我还优化了 UI 的美观度和便捷度。</p><p>AI 会理解错一些点。例如 <a href="https://github.com/CalvinNeo/OpenBoardGame/commit/ab26a9c00cd5d9720b39bf3e248b672881cb52ed" target="_blank" rel="noopener">https://github.com/CalvinNeo/OpenBoardGame/commit/ab26a9c00cd5d9720b39bf3e248b672881cb52ed</a> 这个修复 commit 就展示了 AI 对璀璨宝石最大数量的多次理解问题：</p><ol><li>一开始，它根本没有实现这个限制。</li><li>后面，它实现为只有超过 10 才不能拿，但是从 9 到 12 这个行为是被它允许的。</li><li>最后，才修改对了。</li></ol><p>AI 会漏掉一些情况。例如 <a href="https://github.com/CalvinNeo/OpenBoardGame/commit/c8e82736975325f0b9300b471524ec86b005129e#diff-794e220aafafcfac193a89abd6fb142d92255d544728f217d64e7a1162c79e28" target="_blank" rel="noopener">https://github.com/CalvinNeo/OpenBoardGame/commit/c8e82736975325f0b9300b471524ec86b005129e#diff-794e220aafafcfac193a89abd6fb142d92255d544728f217d64e7a1162c79e28</a> 这个修复 commit 展示了 AI 对用户离开规则的遗漏：</p><ol><li>先前，AI 处理了 in game 的情况。此时如果最后一个活人玩家离开游戏，那么这个房间可以被手动清理掉。</li><li>但是，它漏掉了 in lobby 的情况。此时房间并没有开始游戏，那么玩家离开房间（比如创建一个新的房间）不会导致该房间出于可以被清理的状态。</li></ol><p>经验：</p><ul><li>让 agent 缩小阅读的范围。例如可以告诉它“这是完全的前端修改，你不需要看后端代码或者其他游戏的代码”，这样它就可以更快解决问题，并且能节省不少额度。这随着项目增大是尤为有效的，因为看 thinking 是可以发现它有倾向去学习其他代码是怎么做的。</li><li>让 AI 先整理信息，然后再写一个 design 征求意见非常重要。因为 AI 在实现的时候还是偏向于漏点东西的，这也可能是出于对问题的不正确理解。</li><li>逻辑比较独立的部分，可以让 AI 整理出测试。虽然目前也没看到 AI 会主动修改代码从而 break 掉测试，但这样会更有自信。</li></ul><h1 id="Jan-30-Jan-31"><a href="#Jan-30-Jan-31" class="headerlink" title="Jan 30 - Jan 31"></a>Jan 30 - Jan 31</h1><p>这几天主要实现了出包魔法师、猜狐狸、截码战三个游戏。</p><h1 id="Feb-1-Feb-3"><a href="#Feb-1-Feb-3" class="headerlink" title="Feb 1 - Feb 3"></a>Feb 1 - Feb 3</h1><p>这几天主要实现了角斗士棋、Store&amp;Load、印象花语、AI 画物语四个游戏。主要是由 Gemini 生成游戏说明书，再由 Codex 生成 design。等我 Review 了之后，再实现代码。</p><p>角斗士棋和印象花语中都涉及到了拖拽旋转对象的设计。我发现 AI 在适配手机端上的拖拽是相对比较蠢的，需要手动告诉怎么搞。</p><p>角斗士棋实现起来很简单，但是其实要打磨的很多：</p><ul><li>手机版如何精准地把方块放到格子里面就很难困难。因为每个方块的焦点并不是方块的重心。后来加了上下左右移动的键，但是感觉还是很蠢。</li><li>还有一些奇怪的 bug。</li></ul><p>印象花语这个游戏，我对实体版的一些物理效果进行了模拟：</p><ul><li>因为印章可以印得深浅不一，所以我增加了点按时长来决定透明度的机制。并且我也增加了根据色彩学原理来模拟印章叠加的效果。</li><li>因为可以倾斜印章，使得只盖一半。所以我增加了一个 Mask 机制，可以去 Clip 印章。这个就好像我在按之前垫了一张废纸，这样章可以只盖一半在画布上。</li></ul><p>我觉得这一部分涉及到物理世界的原理的，应该只有人类能意识到并且提出来。</p><p>Store&amp;Load 是非常有意思的点，AI 当时给我提供了一个完全不一样的解决方案。例如它执着于解决 Session 的问题，如何处理同名玩家和同名 Room。但是我提出了一个非常简单的做法，就是完全新建一个 Room，然后新玩家可以进去 Claim 老玩家的 Seat。这样做能够减少很多的复杂度。另外，我还发现新创建的房间并不支持 Auto Save 选择了，这个我觉得也是 AI 本身缺乏引申的能力。</p><p>AI 画物语的实现是跟其他桌游完全不一样的，因为要涉及生成图。所以我开了 <a href="https://github.com/CalvinNeo/aidixitgen" target="_blank" rel="noopener">https://github.com/CalvinNeo/aidixitgen</a> 这个 repo。</p><p>一开始它是用的 pollinations 做的。然后我发现这个工具经常生成水母，然后问了他，他说是提示词的问题：</p><blockquote><p>为什么这个代码里面的 prompt 这么喜欢生成 jellyfish？<br>这是一个非常经典且有趣的 AI 生成现象，通常被称为 “模式坍塌” (Mode Collapse) 或 “安全区偏差”。出现大量水母（Jellyfish）、热气球、云朵、漂浮的岛屿，原因主要有三点：<br>关键词的“刻板印象”： 你的代码里用了 surreal（超现实）、dreamlike（梦幻）、whimsical（异想天开）这些词。 在 AI 的训练数据（Latent Space）里，水母是“梦幻”和“超现实”的超级代表物。它们透明、发光、漂浮，长得就像外星生物，完美符合 AI 对“梦境”的理解。只要你还要“唯美”和“超现实”，AI 就会懒惰地选择水母。<br>Prompt 缺乏约束： 我们在请求 AI 写创意时，给的指令太宽泛了（”Generate ONE single short… description”）。当指令太宽泛时，AI 会倾向于输出它认为“最稳妥、最符合题意”的高概率词汇。对于 Dixit 风格，水母就是那个“最稳妥”的答案。<br>风格修饰词的引导： 我们在画图时强制加了 soft colors（柔和色彩）和 floating（虽然没写但 dreamlike 隐含了漂浮感）。这也极大地增加了生成海洋生物或云端生物的概率。</p></blockquote><p>然后，我又发现它生成的东西太容易重复了，例如几乎生成了两个只是风格不同的钟。所以我后面又设计了双主体的方案，结果最后生成的结果是比较理想的了。</p><p>后来发现 pollinations 一直 502 了，所以就换 hugglingface 了。</p><h1 id="Feb-4-Feb-6"><a href="#Feb-4-Feb-6" class="headerlink" title="Feb 4 - Feb 6"></a>Feb 4 - Feb 6</h1><p>这几天主要实现了前端美化、Flip 7、德国心脏病、绝妙误解。</p><p>主要是由 Gemini 生成游戏说明书以及设计。然后由 Codex 去实现。但是我要求 Codex 在实现前先问我不清楚的项目，而不是自己随便实现一版。事实发现，让 Codex 去问一下自己不知道的，而不代替我做决定，是很重要的。</p><p>AI 前端的主要问题：</p><ul><li>UI 直白<br>  例如用一个列表表示玩家信息。用一个表格表示当前状态。这对玩家而言体感不好，感觉是在上班。</li><li>没有交互设计<br>  特别是手机端玩家，操作的时候需要翻来翻去。</li><li>UI 可能存在 Bug<br>  例如在手机端会发现 Flip 7 的 Game 面板会非常少。</li></ul><p>德国心脏病的开发是非常典型的。主要包含几点：</p><ul><li>得到的开发计划是经典版的 Halli Galli，也就是牌上的水果一定是相同的。这个跟我们玩的不一样，所以后面让它开发了一个 DLC 一样的东西。</li><li>电脑根本没有给翻拍和按铃的等待时间，所以加上 bot 之后，基本 bot 都是秒按铃，秒翻牌，根本没法玩。即使没有 bot，我们也需要考虑人类的反应时间，以及各个网络的延迟。所以我这里要求加了等待 3s 的按铃时间，以及在点击翻拍后，有一个 1s 的倒计时，方便大家准备看新水果。</li><li>AI 生成的界面依然是列表，这个我让 AI 改成了围成一个圆。</li></ul><h1 id="Feb-7-Feb-10"><a href="#Feb-7-Feb-10" class="headerlink" title="Feb 7 - Feb 10"></a>Feb 7 - Feb 10</h1><p>主要工作是修复之前各个游戏的问题。例如：</p><ul><li>cabo 的规则实现错误</li><li>flip7 的相关问题</li><li>继续修理绝妙误解中糟糕的前端问题</li><li>印章话语支持投票</li><li>支持日志复制</li><li>发现之前代码中有 valid 检查，一直是死代码，启用</li></ul><p>实现了新游戏，包括：Gold Rush、Hanabi、Cyber Pictures、The Gang、逻辑猫。</p><p>对 FRONTEND.md 进行了修订，后续的移动端 UI 开始要好一些了。</p><p>此外，还支持了新的功能：</p><ul><li>Download Memories，用来下载整个游戏的回放<br>  实现的过程中，发现设计非常死板。例如花了大篇幅去展示什么 user id，但是游戏内容就缩在一小列里面。而且他非常喜欢表格。</li><li>在游戏过程中开启 Auto Save</li></ul><h1 id="Feb-11-Feb-12-Feb-26-27"><a href="#Feb-11-Feb-12-Feb-26-27" class="headerlink" title="Feb 11 - Feb 12, Feb 26-27"></a>Feb 11 - Feb 12, Feb 26-27</h1><p>主要支持了我画我猜、方鸟、小早川、牛头王。快艇骰子。</p><p>方鸟这个游戏非常有趣，是 Gemini 生成了一段很有歧义的规则说明，然后 Codex 直接按照错误的方式理解了。</p><h1 id="Feb-28-Mar-12"><a href="#Feb-28-Mar-12" class="headerlink" title="Feb 28 - Mar 12"></a>Feb 28 - Mar 12</h1><p>这段时间开始尝试探索 AI 多模态能力的边界。主要开发的游戏有历史奇旅、卡卡颂、德州扑克、project L、璀璨宝石宝可梦。此外，还尝试了诸如 Claude code 等其他模型的能力。</p><p>其中：</p><ul><li>历史奇旅主要是识别质量较高的游戏卡牌</li><li>卡卡颂是识别卡牌，并能拓扑地利用它们<br>  实际上，识别是最难的。我大部分时间用来调试 Codex 生成 72 个基础地图。Codex 经常是分不清城堡是否占边，是否占角。还分不清城堡和道路的重合关系。在生成 svg 的时候，非常喜欢用圆而不是二次曲线去拟合，导致生成的图片很僵硬。<br>  但是，一旦生成成功，后面包括地图连通性判断就非常简单，直接一遍过。</li><li>project L 是阅读游戏说明书，然后生成<br>  这里面主要是前端显示问题。就是一个 svg 生成好好的，让它转成前端，就变成了长方形的长条而不是正方形的格子了。</li><li>宝可梦是识别质量较低的游戏卡牌<br>  可以说非常痛苦，目前来看人工校对是必要的。强烈建议对于这种任务，让 ai 首先生成一个校对工具，这样每次的错误至少是能收敛的。另外，确实是有 corner case AI 无法识别，这样就只能人工标注。<br>  另外，这种 case，AI 喜欢写一段识别代码，我建议提示 AI 不要用 tesserocr。另外，这种类似于“蒸馏”的方式其实挺好。</li></ul><p>此外，还进行了一些优化：</p><ul><li>将 app.js 拆分以减少 tokens 开销</li><li>引入 vendor 减少 cdn 故障导致的无法游戏</li><li>增强游戏选择界面</li><li>引入帮助功能</li></ul>]]></content>
    
    
    <summary type="html">&lt;p&gt;作为一个桌游爱好者，我打算用 Codex 去 Vibe 一个桌游模拟器，这样我可以尝试自定义规则和 Bot 强度。这篇文章我会介绍我 Vibe 的经验。&lt;/p&gt;
&lt;p&gt;我的项目是 &lt;a href=&quot;https://github.com/CalvinNeo/OpenBoardGame%E3%80%82&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://github.com/CalvinNeo/OpenBoardGame。&lt;/a&gt;&lt;/p&gt;</summary>
    
    
    
    
    <category term="数据结构" scheme="http://www.calvinneo.com/tags/数据结构/"/>
    
    <category term="agent" scheme="http://www.calvinneo.com/tags/agent/"/>
    
  </entry>
  
  <entry>
    <title>3FS 学习</title>
    <link href="http://www.calvinneo.com/2026/01/15/study-3fs/"/>
    <id>http://www.calvinneo.com/2026/01/15/study-3fs/</id>
    <published>2026-01-15T15:09:06.000Z</published>
    <updated>2026-05-27T13:26:57.728Z</updated>
    
    <content type="html"><![CDATA[<p>学习 3FS。</p><a id="more"></a><h1 id="原理"><a href="#原理" class="headerlink" title="原理"></a>原理</h1><h2 id="解决什么问题"><a href="#解决什么问题" class="headerlink" title="解决什么问题"></a>解决什么问题</h2><p>3FS 作为一个训推一体框架，解决的几个痛点是：</p><ul><li>痛点一：AI 训练中海量数据的“随机读取”瓶颈 在模型训练时，数据加载器（Dataloader）需要不断对海量的训练样本进行全局 Shuffle（打乱重排）以防止模型过拟合。这会产生极其庞大的纯随机读取操作。在传统架构下，为了不让 GPU 等待数据，往往需要预先将数据搬运到计算节点的本地硬盘上。3FS 聚合了全局极高的并发随机读取能力，让计算节点可以直接像访问本地文件一样，在全局数据集上进行无感知的随机采样，省去了繁琐的数据预取和搬运。</li><li>痛点二：传统文件系统“读缓存失效”造成的内存浪费 主流的操作系统高度依赖 Page Cache（页缓存）和预读机制来加速文件读取。但在 AI 训练的随机洗牌场景下，读过的数据短时间内几乎不会再被重复读取。这意味着传统的“缓存”不仅完全无效，还会大量无谓地侵占计算节点的系统内存，甚至干扰训练任务的稳定性。3FS 针对这一点进行了彻底剥离，它主要依赖 Linux AIO 和 io_uring 直接处理 I/O 操作（Direct I/O），完全抛弃了对本地文件缓存的依赖，把宝贵的内存还给了计算任务。<br>  关于这一点，可以参考 Are You Sure You Want to Use MMAP in Your Database Management System？ 这个文章。我们在使用 tantivy 的时候，也发现 tantivy 的 mmap 会逃逸 tiflash 读节点的内存管理，导致频繁出现 memory over limit 的软限制，导致查询报错。<br>  我们开发的 tici 使用的 tantivy 会使用 mmap，并且会对每个 fragment 做 cache。mmap 无法直接控制，frag cache 目前也是必须的，所以这一块的内存管理很被动。</li><li>痛点三：超大规模集群的“并行 Checkpoint”风暴 万卡集群在训练千亿/万亿参数大模型时，需要高频保存模型权重状态（Checkpoint）以防硬件故障导致训练白费。这就要求系统必须在极短的时间内，将几十 TB 甚至上百 TB 的数据同时且高并发地砸进磁盘。3FS 极高的吞吐上限（官方测试可达极高的 TiB/s 级别），将写入 Checkpoint 导致的集群全量停机等待时间压缩到了极致。<br>  这里面的 Checkpoint 不仅是模型参数 Weights（每一层的权重），还包括优化器的状态，如 Adam 优化器的一阶矩和二阶矩，梯度和一些元数据等。主要的内容是 Adam 部分。<br>  假设你在训练一个 1,750 亿参数（175B）的模型，以半精度（16-bit）保存，权重约 350GB。但加上优化器状态后，一个完整的 Checkpoint 可能高达 2TB 左右。如果是一个万亿参数（MoE 架构）的模型，单个 Checkpoint 达到 10TB-20TB 并不罕见。为了容灾和实验回溯，集群通常会保留多个步数的副本，因此总占用达到百 TB 级别是非常正常的。</li><li>痛点四：推理阶段 KV Cache 的容量与成本限制 在处理超长上下文（Context）的 LLM 推理时，保存在昂贵 GPU 显存或系统内存（DRAM）中的 KV Cache 会迅速触及容量天花板。3FS 极低的读写延迟和超高并发，让 DeepSeek 能够实现“将 KV Cache 卸载到廉价高速磁盘（SSD）上”的技术方案。这为超长文本的推理提供了一种极具性价比的替代路径。</li></ul><h2 id="3FS-的方案"><a href="#3FS-的方案" class="headerlink" title="3FS 的方案"></a>3FS 的方案</h2><p>对此，3FS 的方案是：</p><ol><li>把显存 offload 到 SSD 上。主要还是借助 NVMe 和 RDMA 技术。</li><li>Global Prefix Caching：我理解这一部分主要的优化点是超级冗长的系统提示词 (System Prompt) 以及各种套娃。</li></ol><p>另外，是 KVCache 的内存开销，其实是和上下文线性相关的<br><img src="/img/3fs/kvcache-complexity.png"></p><h1 id="设计-–-design-notes"><a href="#设计-–-design-notes" class="headerlink" title="设计 – design notes"></a>设计 – design notes</h1><p>3FS 中有个 design notes，这里主要是概括了 design notes。</p><h2 id="需要解决的问题"><a href="#需要解决的问题" class="headerlink" title="需要解决的问题"></a>需要解决的问题</h2><p>OSS 方面：</p><ul><li>现在的 OSS 并不支持原子地移动一系列文件或者一整个目录，或者递归地删除整个目录。而 3FS 的场景（实际上数据库的场景也是这样）涉及要创建一个临时目录，然后对这个目录写入数据，最后将这个目录 move 到最终的位置。</li><li>3FS 需要广泛使用 symbolic 或者 hard 链接</li><li>提供一个熟悉的文件接口，我理解这也是 3FS 选择 FUSE 的原因</li></ul><p>FUSE 方面：</p><ul><li>在 <a href="/2025/03/09/learn-fuse/">Fuse 学习</a>一文中介绍了为什么 FUSE 不支持 Zero Copy。但这也是 FUSE 的缺点之一。</li><li>FUSE 使用一个由 spin lock 保护的多线程共享的队列。3FS 团队的测试显示，400K 4KiB reads per second 的写入负载之后，因为 lock contention 的方法。</li><li>Linux 5.x 上的 fuse 不支持对一个文件的并发写。所以很多需要更大带宽的程序会并发写多个文件。</li><li>对小的随机非对齐读性能不好，SSD 和 RDMA 网络的带宽没有被利用充分。</li></ul><p>但是将 client 实现为一个 VFS 内核模块则能解决上面说的问题，但也更有挑战性。内核的 bug 更难定位和修复。此外，升级的时候需要停掉所有访问这个 fs 的进程，或者重启。</p><p>因此，3FS 选择在 FUSE daemon 里面设计一套原生的 client，由它来支持异步的 Zero copy IO。其中，File meta operation 例如 open、close 等还是被 FUSE daemon 处理。但是在 open 的时候会把拿到的 fd 通过 native API 注册。然后就可以通过 native client 去读取数据了。</p><p>这个 API 类似于 io_uring，其中关键结构如下：</p><ul><li>Iov<br>  user process 和 native client 共享的内存</li><li>Ior<br>  user process 和 native client 通过这个 ring buffer 进行交互。具体方式类似于 io_uring。<br>  请求会被按照 io_depth 攒批执行，不同的 batch 的执行是并行的。</li></ul><h2 id="Metadata-存储"><a href="#Metadata-存储" class="headerlink" title="Metadata 存储"></a>Metadata 存储</h2><h3 id="chunk-的分布"><a href="#chunk-的分布" class="headerlink" title="chunk 的分布"></a>chunk 的分布</h3><p>文件按 chunk 为粒度，被条带化到多个 replication chain 中。<br>创建新文件的时候，会根据 stripe size，使用 round robin 的方式，选择一系列的 chain。选出来的这些 chain，会随机分给不同的 chunk 写入。</p><h3 id="存储-file-attributes"><a href="#存储-file-attributes" class="headerlink" title="存储 file attributes"></a>存储 file attributes</h3><p>为什么 3FS 的 inode 里的 length 会不准确？因为写路径走的是 CRAQ，而 inode 是在 metadata service 里面的。如果每次写操作完都更新一下 metadata，那么会多一次 metadata RTT，写放大严重，吞吐和延迟都会变差。<br>但是如果 metadata 迟迟不更新，例如是 100mb，而客户端写到 120mb 就挂了，此时，虽然数据已经通过 CRAQ 持久化到 chunk 存储了，但因为读的时候从 metadata 获得的长度偏小，所以还是在效果上丢失数据。<br>一种方式是按照 interval 上报更新 metadata，但这就存在不一致窗口。不过先考虑容灾问题，大概有两个方案：</p><ol><li>重启之后，从 chunk 存储中恢复数据，并由此更新 metadata。但是从 chunk 扫描数据恢复的代价很大</li><li>3FS 的设计是由 client 按照 interval 上报 max writer position，因为 client 上报的 position 一定是已经被 tail 提交了的，所以是安全的。但是如果 client 长期丢失，那么 gap 就得通过第一种方式补齐了。</li></ol><h2 id="Chunk-存储"><a href="#Chunk-存储" class="headerlink" title="Chunk 存储"></a>Chunk 存储</h2><p>Suppose there are 6 nodes: A, B, C, D, E, F. Each node has 1 SSD. Create 5 storage targets on each SSD: 1, 2, … 5. Then there are 30 targets in total: A1, A2, A3, …, F5. If each chunk has 3 replicas, a chain table is constructed as follows.</p><table><thead><tr><th align="center">Chain</th><th align="center">Version</th><th align="center">Target 1 (head)</th><th align="center">Target 2</th><th align="center">Target 3 (tail)</th></tr></thead><tbody><tr><td align="center">1</td><td align="center">1</td><td align="center"><code>A1</code></td><td align="center"><code>B1</code></td><td align="center"><code>C1</code></td></tr><tr><td align="center">2</td><td align="center">1</td><td align="center"><code>D1</code></td><td align="center"><code>E1</code></td><td align="center"><code>F1</code></td></tr><tr><td align="center">3</td><td align="center">1</td><td align="center"><code>A2</code></td><td align="center"><code>B2</code></td><td align="center"><code>C2</code></td></tr><tr><td align="center">4</td><td align="center">1</td><td align="center"><code>D2</code></td><td align="center"><code>E2</code></td><td align="center"><code>F2</code></td></tr><tr><td align="center">5</td><td align="center">1</td><td align="center"><code>A3</code></td><td align="center"><code>B3</code></td><td align="center"><code>C3</code></td></tr><tr><td align="center">6</td><td align="center">1</td><td align="center"><code>D3</code></td><td align="center"><code>E3</code></td><td align="center"><code>F3</code></td></tr><tr><td align="center">7</td><td align="center">1</td><td align="center"><code>A4</code></td><td align="center"><code>B4</code></td><td align="center"><code>C4</code></td></tr><tr><td align="center">8</td><td align="center">1</td><td align="center"><code>D4</code></td><td align="center"><code>E4</code></td><td align="center"><code>F4</code></td></tr><tr><td align="center">9</td><td align="center">1</td><td align="center"><code>A5</code></td><td align="center"><code>B5</code></td><td align="center"><code>C5</code></td></tr><tr><td align="center">10</td><td align="center">1</td><td align="center"><code>D5</code></td><td align="center"><code>E5</code></td><td align="center"><code>F5</code></td></tr></tbody></table><p>这里的 Version 是配置的 Version，节点下线会导致这个增大。</p><p>这里的一个 Chain 类似于一个 Raft Group 的概念。但是它也不是像 TiKV 一样跟某一段数据绑定的。一个 Chain 可以被多个 chain table 包含。引入 chain table 的概念，这样对于每一个 file，metadata service 就可以为它选一个 chain table，并根据这个 table 中的 Chain 去 strip 这个 file 的所有 chunk。</p><h3 id="Balanced-traffic-during-recovery"><a href="#Balanced-traffic-during-recovery" class="headerlink" title="Balanced traffic during recovery"></a>Balanced traffic during recovery</h3><p>如果一个节点 A 故障了，就需要由 Chain 中的其他节点来承担原来 A 的流量。而之前的 Chain table 中，A 节点基本上只和 B、C 玩。</p><p>在新的架构中，A 在 Chain 2 里和 B/D 在一起，在 Chain 5 里和 C/F 在一起。</p><h3 id="Data-replication"><a href="#Data-replication" class="headerlink" title="Data replication"></a>Data replication</h3><p>一个 Write request 可能是从 client 或者 Chain 的前驱发送出来的。一个节点收到 Write request 后的处理：</p><ul><li>校验 write request 中的 chain version。</li><li>通过 RDMA Read 去 pull 写入的数据。如果 client 或者前驱挂掉了，导致拿不到数据。写入就 abort。</li><li>Once the write data is fetched into local memory buffer, a lock for the chunk to be updated is acquired from a lock manager. Concurrent writes to the same chunk are blocked. All writes are serialized at the head target.</li><li>读取这个 Chunk 的 committed version，对它 apply change，然后将更新后的版本存储为 pending version。版本号是单调连续递增的。</li><li>If the service is the tail, the committed version is atomically replaced by the pending version and an acknowledgment message is sent to the predecessor. Otherwise, the write request is forwarded to the successor. When the committed version is updated, the current chain version is stored as a field in the chunk metadata.</li><li>When an acknowledgment message arrives at a storage service, the service replaces the committed version with the pending version and continues to propagate the message to its predecessor. The local chunk lock is then released.</li></ul><h1 id="实现"><a href="#实现" class="headerlink" title="实现"></a>实现</h1><h2 id="Chain-replication"><a href="#Chain-replication" class="headerlink" title="Chain replication"></a>Chain replication</h2><p>CRAQ 的最大的特点是：复制链路是固定的，而 Raft 的复制链路因为存在 Quorum 是随机的。因此，对于单个 entry 的复制，Raft 的延迟可能会比 CRAQ好，但是 CRAQ 的延迟和吞吐更稳定，可以被预测。</p><p>延迟维度的对比：</p><ol><li>Raft 最优<br> a. 条件：follower 延迟几乎一致、没有网络问题<br> b. commit latency 约等于 RTT</li><li>Raft 最劣<br> a. 条件：quorum 中某个 follower 抖动，例如出现了 Write Stall<br> b. Throughput 约等于 1 / avg([max(follower latency)]，这里的 avg 因为抖动是不均匀的，因为多数派每次都不一定一样。</li><li>CRAQ 最优<br> a. 条件：链上每个 node 的 hop_latency 相同，管道用满<br> b. Throughput ≈ min(hop throughput)，这个很简单，木桶效应嘛。</li><li>CRAQ 最劣<br> a. 条件：链上每个 node 变慢<br> b. 存在木桶效应，每个 entry 的提交都被固定变慢。</li></ol><p>吞吐维度的对比：</p><ol><li>Raft 最优<br> a. 条件：同延迟。<br> b. Throughput ≈ follower replication rate</li><li>Raft 最劣<br> a. 条件：同延迟。<br> b. commit latency 等于 max(quorum follower latency)，出现了木桶效应</li><li>CRAQ 最优<br> a. 条件：链上每个 node 的 hop_latency 相同<br> b. commit latency 约等于 chain_length * hop_latency。因为 CRAQ 在 Tail 返回确认，所以 3 副本也是两次数据传输，和 Raft 的 1 个 RTT 是接近的。但是链长了，CRAQ 就会变慢。</li><li>CRAQ 最劣<br> a. 条件：存在一个很慢的 node<br> b. 同上，但是木桶效应被放大很明显。</li></ol><p>CRAQ 的其他特点：</p><ol><li>没有 quorum 容错</li><li>failover 不需要重新选主，但需要重构链。需要根据挂的是哪一个来讨论。</li></ol><h1 id="Reference"><a href="#Reference" class="headerlink" title="Reference"></a>Reference</h1><ul><li><a href="https://github.com/deepseek-ai/3FS/blob/main/docs/design_notes.md" target="_blank" rel="noopener">https://github.com/deepseek-ai/3FS/blob/main/docs/design_notes.md</a></li></ul>]]></content>
    
    
    <summary type="html">&lt;p&gt;学习 3FS。&lt;/p&gt;</summary>
    
    
    
    
    <category term="数据库" scheme="http://www.calvinneo.com/tags/数据库/"/>
    
    <category term="aiinfra" scheme="http://www.calvinneo.com/tags/aiinfra/"/>
    
  </entry>
  
  <entry>
    <title>用英文写作计算机博客</title>
    <link href="http://www.calvinneo.com/2026/01/12/write-blogs-in-english/"/>
    <id>http://www.calvinneo.com/2026/01/12/write-blogs-in-english/</id>
    <published>2026-01-12T14:42:32.000Z</published>
    <updated>2026-01-18T19:06:49.071Z</updated>
    
    <content type="html"><![CDATA[<p>介绍下用英文写作计算机博客的一些经验。</p><a id="more"></a><h1 id="常见的表达"><a href="#常见的表达" class="headerlink" title="常见的表达"></a>常见的表达</h1><h2 id="避免直接翻译汉语词"><a href="#避免直接翻译汉语词" class="headerlink" title="避免直接翻译汉语词"></a>避免直接翻译汉语词</h2><p>少用汉语式的名词化表达，例如“执行 xx 行动”，“处理 xx 过程”。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">the lock is released</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">the release of the lock is performed</span><br></pre></td></tr></table></figure><p>不要自以为是去形容词化。如下，adaptive 的意思是自适应的，和我们实际要指代的“适配代码的行为”是不对应的。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">We need to **adapt** the code to the new behavior. However, this **adaptive** work is not easy.</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">... However, the adaptation work is not easy.</span><br></pre></td></tr></table></figure><h3 id="实词虚化、具体词抽象化"><a href="#实词虚化、具体词抽象化" class="headerlink" title="实词虚化、具体词抽象化"></a>实词虚化、具体词抽象化</h3><p>汉语中，很多虚词功能是通过复用实词来的。但是在英语中，如果有对应的虚词，就不要直接翻译汉语中的实词了。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br></pre></td><td class="code"><pre><span class="line">Based on this reason, ...</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">As a result, ...</span><br><span class="line">For this reason, ...</span><br></pre></td></tr></table></figure><p>又例如下面的“带来麻烦”，这里的带来没必要用 bring 这个实词</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br></pre></td><td class="code"><pre><span class="line">brings lots of trouble to investigate ...</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">made the issue difficult to investigate ...</span><br><span class="line">// Or</span><br><span class="line">significantly complicated the investigation to ...</span><br></pre></td></tr></table></figure><p>又例如下面的“定位到问题”</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">The problem is located by ...</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">The problem was also detected by ...</span><br></pre></td></tr></table></figure><h3 id="虚词实化"><a href="#虚词实化" class="headerlink" title="虚词实化"></a>虚词实化</h3><p>但是，一些汉语中的虚词，在英语中要实化。例如，“这会导致难以理解的代码”，就是</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">It could result in confusing codes.</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">It could generate confusing codes.</span><br></pre></td></tr></table></figure><p>这个原则甚至不限于词，对于任何表达都是这样。<code>This may cause problems</code>，建议直接具体一点说是什么 problems。如</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br></pre></td><td class="code"><pre><span class="line">This behavior can cause a deadlock when two tasks hold the mutex across an await point.</span><br><span class="line">// Or</span><br><span class="line">This design is problematic in terms of concurrency safety, because it allows a task to hold a mutex across an await point.</span><br></pre></td></tr></table></figure><h2 id="需要调整句子结构"><a href="#需要调整句子结构" class="headerlink" title="需要调整句子结构"></a>需要调整句子结构</h2><p>尽量避免使用形式主语 It is 或者 there is 等来拖长表达</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">It is hard to inspect the state.</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">Inspecting the state is difficult.</span><br></pre></td></tr></table></figure><p>但注意，非形式主语是可以用 it 的，如下所示，这好过说 “I think” 等</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br></pre></td><td class="code"><pre><span class="line">It indicates ...</span><br><span class="line">This implies ...</span><br><span class="line">The result shows ...</span><br></pre></td></tr></table></figure><p>从下面的例子中，能感觉到动名词前置的用法，更干练</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">I don&apos;t think the service itself should persist the updated configuration to the config file. Instead, this should be handled by the operator.</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">Persisting the updated configuration should be the responsibility of the operator, not the service itself.</span><br></pre></td></tr></table></figure><h2 id="语序"><a href="#语序" class="headerlink" title="语序"></a>语序</h2><p>副词顺序应该稳定在动词后。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">We must immediately do this.</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">We must do this immediately.</span><br></pre></td></tr></table></figure><h2 id="其他"><a href="#其他" class="headerlink" title="其他"></a>其他</h2><p>下面的表达，相比更能体现出不仅不能做之前说的，也不能做现在说的。而 Also 显得更像一个连接。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br></pre></td><td class="code"><pre><span class="line">Also, A can&apos;t do sth.</span><br><span class="line"></span><br><span class="line">===&gt;</span><br><span class="line"></span><br><span class="line">Moreover, A can not do sth either.</span><br><span class="line">We can&apos;t ..., nor can we ... .</span><br></pre></td></tr></table></figure><h1 id="使用更恰当的单词"><a href="#使用更恰当的单词" class="headerlink" title="使用更恰当的单词"></a>使用更恰当的单词</h1><h2 id="使用更精确的词"><a href="#使用更精确的词" class="headerlink" title="使用更精确的词"></a>使用更精确的词</h2><p>避免使用 </p><p>possible -&gt; feasible</p><h2 id="使用恰当的搭配"><a href="#使用恰当的搭配" class="headerlink" title="使用恰当的搭配"></a>使用恰当的搭配</h2><p>动名词搭配。</p><h2 id="辩证具体含义的差别"><a href="#辩证具体含义的差别" class="headerlink" title="辩证具体含义的差别"></a>辩证具体含义的差别</h2><p>如：</p><ul><li><a href="https://english.stackexchange.com/questions/277073/which-is-correct-confident-in-or-confident-of?newreg=42e56658dfef4cb1b187d36e10d24f6d" target="_blank" rel="noopener">be confident of 和 be confident in</a></li></ul><h1 id="宏观写法"><a href="#宏观写法" class="headerlink" title="宏观写法"></a>宏观写法</h1><p>先给结论，再给解释（Top-down writing）。</p><h1 id="特定场景的表达"><a href="#特定场景的表达" class="headerlink" title="特定场景的表达"></a>特定场景的表达</h1><h2 id="偏向数据分析"><a href="#偏向数据分析" class="headerlink" title="偏向数据分析"></a>偏向数据分析</h2><p>表达“A使用的内存占A的调用者的比重，相比B使用的内存占B调用者的比重是相近了”，从书面到口语。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br></pre></td><td class="code"><pre><span class="line">// 这里的 footprint 更专业术语一点。</span><br><span class="line">A’s memory footprint relative to its callers is similar to B’s.</span><br><span class="line"></span><br><span class="line">A and B exhibit similar memory-to-caller ratios.</span><br><span class="line"></span><br><span class="line">The proportion of memory used by A relative to its callers is similar to that of B.</span><br><span class="line"></span><br><span class="line">A and B have similar memory usage ratios with respect to their callers.</span><br></pre></td></tr></table></figure>]]></content>
    
    
    <summary type="html">&lt;p&gt;介绍下用英文写作计算机博客的一些经验。&lt;/p&gt;</summary>
    
    
    
    
    <category term="English" scheme="http://www.calvinneo.com/tags/English/"/>
    
  </entry>
  
  <entry>
    <title>太原游记</title>
    <link href="http://www.calvinneo.com/2026/01/04/traval-in-taiyuan/"/>
    <id>http://www.calvinneo.com/2026/01/04/traval-in-taiyuan/</id>
    <published>2026-01-03T18:20:13.000Z</published>
    <updated>2026-01-09T07:37:06.530Z</updated>
    
    <content type="html"><![CDATA[<p>总的来说，太原是类似于西安的城市，类似的气候，类似的历史文化。但在游玩体验来讲，太原总体略优于西安，主要是：</p><ul><li>西安商业化太浓重了，例如大唐不夜城、城墙上骑车等。当然这也不是坏事，但如果玩的多了，就会觉得商业化严重的景点如同预制菜一样，不能说不好吃，但感觉容易腻</li><li>西安人太多了</li></ul><p>但太原的问题主要是：</p><ul><li>交通很不方便。虽然也看到它有专门的旅游线路，并且感觉是用心的，但体验上确实还有欠缺。</li></ul><a id="more"></a><h1 id="D1"><a href="#D1" class="headerlink" title="D1"></a>D1</h1><p>我们是早上 7.55 的飞机，所以定了个 5.10 的送机，相比打车要贵一点，但好在不需要担心到时候没有车。正好前一天是跨年夜，我们就去雨花台吃了点饭，逛了下雨花万象，就住在雨花了。因为我经典的失眠，导致那一天我实际上没怎么睡。</p><p>在机场躺大座睡了一个小时，然后登机门开了，风吹进来冷的要死，没法继续睡了。好在也快登机了。这鬼飞机是摆渡车，真的冷死，外面还下着雪。</p><p>到了机场，出门就能打车，打车秒打，司机就在停车场，很快就上了车。太原机场离市区是真的近，离太原南站也很近，所以是非常适合中转的城市，这很类似之前去过的张家界。</p><p>北齐壁画博物馆分为三个展厅：</p><ul><li>第一展厅是娄叡墓的壁画。比较有印象的一个是狩猎图，一个是升天图。有印象的细节一个是马一边跑一边被吓得拉屎。另一个是雷公，长得特别奇怪。</li><li>第二展厅徐显秀墓。因为整个博物馆实际上就是盖在这个墓上面的，所以比较有趣。我们可以看到当时发现这个墓的土丘上的盗洞。值得一提的是，虽然我们没法下到最下面的墓里面去，但是旁边有个清晰度贼高的 vr 可以看，不仅壁画很清楚，而且能看到总共五个盗洞。我想，为什么我们看不到棺椁，可能就是这么多盗洞把它们都偷走了了吧。另外，后面在山西博物院中，我们能看到这个墓里面的壁画，但不知道是仿制品，还是移过来的。</li><li>第三展厅是拼好展，里面展示了不同的墓的壁画。</li></ul><p>从北齐壁画博物馆出来，打车非常难打。索性看到有个公交车，就上去，发现好像是个旅游专线，只到双塔公园和火车站。用微信乘车码就可以刷码，坐上去之后，司机一个人发了一张门票一样的东西，实际上就是旅游专线的车票，感觉挺用心的。</p><p>考虑了下，我们觉得先回宾馆比较好。一来，双塔公园离着纯阳宫和宾馆都很远，后面两者反而比较近。二来，我那个包确实太重了。于是打了辆车，直奔宾馆。</p><p>到了宾馆，准备出门，发现山西博物院突然放了一千多张票。于是打电话问，答复说只要是今天的时间预约的就可以来。接线的是个女人，用非常热情的预期表达了对我们的欢迎，非常点赞。所以，我们临时改了计划，去看山西博物院。</p><p>山西博物院就靠着我们的亚朵，中间隔了自然博物馆和图书馆。走过去的时候，门口就排了个小队了。进去之后，发现博物馆里面全是人。这个博物馆主要是四层，重点的是在 2/3 两层，另外有几个重要展物被放到了第 1 层和第 4 层。山西博物馆主要就是看各种青铜器。四楼有很多主体展厅，比如玉、应县木塔、钱币等。总体感觉这博物馆年代比陕历博要老，基本上全是青铜器。比较有印象的一个是那个猫头鹰，还有就是大蜗牛上骑着小蜗牛，以及博物馆徽标上的大鸟上骑小鸟的。有个叫雁鱼铜灯的文物好像被国博借走了，直接在下面放了个不知道是啥的青铜器掩耳盗铃。</p><p>从山西博物馆出来，快五点了，我们赶快打车去河东颐祥阁。事实证明这决定很正确，因为我们到的时候饭店还没开门。但是我们刚坐下来，一堆人就陆陆续续进来了，然后他们就要排队。我们在饭店点了非常多的东西，在此点评一下：</p><ul><li>涮肚 我老婆说这个麻辣烫很好，因为里面的麻酱没有麻酱味，她能接受。我觉得也不错。</li><li>芥末凉粉 很爽口。</li><li>风葫芦 感觉是炸的一个脆皮球，然后里面是很嫩的鸡蛋。我对象很喜欢，我觉得还行，但没有特别很好吃。</li><li>黄米排骨 这个是最贵的，但吃起来感觉一般。</li><li>麻辣串 实际上就是豆干，但是加了他那个口感很椒麻胡辣的汤汁，很好吃。我对象不喜欢胡椒和辣口味，她不怎么喜欢。</li><li>羊腿 我觉得很好吃，我对象也觉得。而且她觉得不太膻。</li><li>槐花 小学里课文学过，但实际上是第一次吃槐花，感觉挺好吃的。没想到是咸的，感觉有点吃野菜的感觉，但并不像野菜那样有茎的感觉。</li><li>野菜丸子 感觉就是北方的那种丸子。</li><li>绿豆糕 挺甜的，但是很好吃。</li><li>烧饼 有点像肉夹馍的馍，但是是很热的，吃起来很香。</li></ul><p>从颐祥阁出来，已经打包了一堆东西，我们打算去鼓楼街再去看看。太原的地铁在长风街这一段特别稀疏，但吃的又都在这里，所以我们走了好远上了地铁。从府西街站下来，我们拐进了帽儿巷。</p><p>刚去的时候，人是一般多，我老婆还说这个离长沙差远了。我们随便走走，就看到一家要排队的卖麻花的，说特别有名。买了黑芝麻和蜂蜜的，感觉就是小时候的味道，有点类似于我奶奶买的那个火腿肠面包。我记得那个面包皮我特别喜欢，就是这个麻花的味道。</p><p>我们后面又买了枣糕和碗托。碗托的那个面感觉挺好吃的，是糯糯的。而很早之前我对象在淘宝买的感觉就很干，跟那种硬胶水一样。不过我当时已经又累又撑，不太吃的下了。</p><p>我准备往地铁站走准备回去了，但是我老婆说，这个鼓楼街，我们还没见到鼓楼呢。于是走到横过来的一条街上，这条街就洋气很多了，都是一些洋人的建筑。再往前走，还有一个巨大的钟楼。这里人就开始特别多了，有种长沙的感觉。</p><p>在认一力吃了几道菜：</p><ul><li>羊肉水饺 点了两种，但感觉都一个味道。我老婆说羊肉味特别大，根本没法吃。我觉得还行，但是感觉很咸。感觉羊肉馅里面加了很多葱姜蒜。另外，第二天我从冰箱里面拿出来的时候，感觉它们散发着一股呕吐味。但加热之后，又能吃了，感觉还行。</li><li>沙棘醪糟 感觉挺好吃的，很清淡，不是特别的酸，也不怎么甜。但是很爽口。</li><li>头脑 头脑里面的羊肉挺好吃的，炖的烂。但是那个粥一样的东西就是一言难尽了。吃起来感觉就好像是熬的很浓的羊油和粥一起煮成的粘稠物。但是呢，它本身又不具备酸甜苦辣咸这些基础的味道，所以吃在嘴里就是一股腥味。</li></ul><p>从认一力出来，人就开始比肩接踵了，于是我们准备往回走了。当时我已经又撑又困到神志不清了，终于走到了柳南地铁站。坐地铁到太原理工大学站，发现到酒店还要走好远。</p><p>晚上回来困得要死，倒头就睡。中途 23 点的时候被老婆叫起来看，原来是酒店对面就放起了烟花。话说山西今年不禁止放烟花，所以大家过年都在放，感觉挺热闹的。汾河上的那条龙今天晚上也是一直亮着灯，还挺好看的。</p><h1 id="D2"><a href="#D2" class="headerlink" title="D2"></a>D2</h1><p>早上起的有点晚，睡了大概 10 个多小时。吃了个酒店的自助餐，顺便把昨天的水饺热了下。</p><p>打车去晋祠，这一路挺远的，司机开得贼快。太原的早晨全是雾霾，路过还有几个大烟囱在冒着白气，可能是供暖工厂吧。到了晋祠，寻思着买了个天龙山的门票。晋祠一进去是非常大的免费公园，里面做了一堆假山雕塑啥的，有个唐太宗的雕塑后面还出现在太原的城市宣传画上，但是感觉雕的人脸都一样。</p><p>晋祠博物馆是挺有意思的，原本我以为它就是一个祠堂，拍照打卡走人这样，但里面挺大的，而且建筑的排布很好看，基本上处处是景。</p><p>游览这个景点，必须要有个讲解。我当时还是用的旅途随身听，感觉讲的是足够了。总而言之，这个晋祠原来是纪念一个儿子的，后来慢慢的，妈妈名气反而更大了，所以改为了主要祭奠妈妈，结果导致儿子的祠堂偏安在整个晋祠的最角落。</p><p>一进门就是晋祠最网红的孙悟空同款了，但其实它并不是祠堂，而是一个戏台，叫水镜台。很多人在正面拍照合影，走到背面有个康熙题的匾额。</p><p>然后，我们就按照旅途随身听的导览逛了，先逛了一遍关帝庙、岳飞庙等，我也认识了一堆歇山顶、硬山顶等屋顶的类型，从而判断规格。然后就到了儿子的唐叔虞的祠堂。这哥们属实惨，真的就在最角落，门口还贼小。要不是随身听，我可能都不会进去。</p><p>从唐叔虞祠堂出来，就看到前面一堆人。这里应该就是晋祠最核心的地方，也就是圣母堂了。圣母堂有几点比较独特：</p><ul><li>堂前的十字桥，据说是首创。这个桥在晋祠外围的免费公园中也有个拙劣的模仿。</li><li>七开间的规格，每个柱子上都有龙盘旋，这些龙长得都不一样，有的还很抽象。</li><li>圣母堂中的宋代雕塑，惟妙惟肖。</li><li>圣母堂旁边有个周朝的柏树，然后它被另一棵柏树架着。</li></ul><p>从圣母堂转了一圈出来，旁边还有一个苗裔娘娘堂，再绕出来，就到了难老泉。这个说是晋阳第一泉，感觉就是比较谦虚了，毕竟无锡还有个天下第二泉呢。这个亭子是北齐建的，所以可以看到斗拱是非常的大。虽然后面在嘉靖年返修了，但仍然是使用了北齐的手法。而它对面的亭子，斗拱就很小，一看就是明清的建筑。其实在晋祠中，分布着不同朝代的建筑，我们都可以从斗拱的大小，以及房顶上鸱吻头和尾巴的比例来分辨年代。晋祠中比较主要的建筑如圣母殿等都是在唐左右建设的。</p><p>难老泉会通过一个龙头流到旁边的一条小溪里面，然后很多人排队在那里接水。我们绕过那群人，就可以走到子乔祠、董寿平美术馆那一带，不过那些就不是重点了。</p><p>我去逛了逛美术馆，我老婆根本就没去逛，而是再绕回来到正面，到了圣母堂之前的一些建筑，如会仙桥、献殿等。比较有意思的是金人台，上面有个非常小的小楼，不知道那是干嘛的。献殿上面有个万历四年的匾额，我对象说是不是那个万历四年春，我说那是庆历四年春。</p><p>从晋祠出来，下一站就是天龙山。但我们必须走很长的路才能穿过外围的免费公园。我们还路过了晋文公艺术博物馆，不过没进去看，不知道里面怎么样。博物馆附近有很多大湖环绕，里面可以划船，但是现在是冬天，这些湖面都结冰了。博物馆前按照先秦的特色，搞了点夯土柱子。绕过博物馆，就到了出口，在这里进行了一些简单的修整，买了个烤红薯和上海阿姨，就去找景区交通。</p><p>公交站排队等了比较长的时间，感觉得有至少十分钟，车来了。因为我们觉得不太妙，所以当时就站的比较靠前，所以有座位坐，但没想到这个公交车是不给站的。所以我们眼睁睁看着后面的人上不来，司机跟他们说等下一班吧。这公交车开了，没一会就上了盘山公路。这路应该太原花了很多钱搞了个旅游公路，很多盘山公路被搞成了类似南浦大桥那样的展线，所谓网红桥。这些展线上停了一堆小轿车，人们在那里往低处拍照片。到了龙门站，我们下车，其实已经花了几十分钟了。</p><p>这天龙山其实很坑，它主要是看石窟的，分为东峰和西峰，但是我们去的时候东峰全封掉了，所以我们就直接往上走回去了，也没再去下面的天龙寺。具体到西峰石窟，里面也是丢的丢，乱涂乱画的也有。后面打了个车去国宝馆，里面看到了东峰石窟中第 8 窟的佛头，说是被日本人之前抢走了的，后来被中国又追回来了。国宝馆旁边还有个数字馆，就是给你看看一些石窟的宣传片。</p><p>从国宝馆直接坐车到晋祠，下了车就打车去植物园。太原植物园很大，一进去是一个湖，湖面已经结冰。围湖种了一圈银杏，叶子很漂亮，不过走近一看发现是假的。因为是冬天，主要能逛的就是几个温室。这几个温室设计的都很好，造景很有一套，并且有高低不同的步道，可以方便游客游览观看不同高度的植物。热带雨林馆里面有个瀑布，大家挤在那里拍照。不过我最喜欢的还是沙生植物馆，一进门是三个大黄柱子，加上远处落日，夕阳照在上面的感觉，给我一种置身于沙漠或者火星表面的感觉。另外还有一个蝴蝶馆，以及一个园艺馆，反正都挺好看的。太原植物园还有萤火虫看，不过冬天就暂停了。</p><p>最终也没等到植物园那个网红天花板亮灯，而是五点二十不到就打车走了。当时打车很容易，不过刚开到市区就发现植物园门口堵红了。我觉得这里的交通设计有问题，汽车要去植物园门口，实际上要走很远到前面掉头才行。这也是太原的一个特点，快速路很多，但是跨越高速路，或者掉头，就很麻烦了。</p><p>去吃东北爱情麻辣拌，在一个非常荒芜的小巷中的破落小平房，里面墙上天花板上都是之前顾客的留言。只能选择加料，以及辣度。我加了麻花和什么的，我对象加了另外一种。这麻辣拌全是素的，全是碳水，要不就是豆制品。旁边一个京爷在吹牛逼说自己老爹是个煤矿里面的啥科长，然后后面北京就不让挖煤了 blabla，说房子贵得要死，幸亏买得早。</p><p>从麻辣烫出来，不远就能走到附近的上帝炸鸡，因为我觉得尽管有外卖，但是没必要一定吃总店。路上还遇到一个卖碗托的路边摊子，不过这次卖的是保德碗托。味道很不一样，我看到她加了一种黄色的酱汁，我想这个应该导致了我们最终吃到的是有点酸味的。加上我们没有要很多辣油，就导致碗托并不腻。不过它的面感觉就和淘宝的荞麦面碗托一样，不如昨天的那个有糯劲。卖碗托的斜对面就是上帝炸鸡，我们点了一份大份的，吃起来感觉就是那种老派炸鸡的感觉，外皮特别脆，但是里面很多汁。不过我觉得得加点辣粉更好吃。</p><p>买完上帝炸鸡，就去利源沾片子。中途路过似乎是太原比较繁华的体育路街区，有个盒马。我们继续往北走就到了沾片子，在一个灯光很暗的小街上。我老婆先去的，然后出来告诉我要排位，留了个电话号码。然后我们就寻思附近转转，这附近可真没啥能转的。一些很拉胯的咖啡厅，泰山啤酒，一堆成人用品店。最后，我们回到那个店，发现人都换了一遍，显然老板看生意好，就没有叫我们。所以我们被迫在店里面站着等位，等的时候，后面最多来了四队人。老板让我们提前点菜，我们点了一堆：</p><ul><li>沾片子 吃起来比较素。我比较喜欢豇豆的那个，感觉保留了蔬菜的本味。</li><li>炒茄子 很好吃，茄子很脆，酱汁很香。</li><li>莜面</li><li>清徐灌肠 实际上也是素的，感觉挺油的</li></ul><p>这家店不光等位慢，做得也慢，出来了就九点多了。打车不太打得到，所以就准备坐地铁，路上还买了一盒草莓，大个的，才 30 块钱。好不容易坐上了地铁，然后下错站了，所幸直接打车算了。不过这样晚上在汾河边散步的计划就泡汤了，实际上那天晚上步道的灯光包括那条龙也没开，所以就算了。</p><h1 id="D3"><a href="#D3" class="headerlink" title="D3"></a>D3</h1><p>早上七点五十就起来了，先把昨天的炸鸡拿到宾馆餐厅里面热热，顺便再吃吃他的刀削面和豆花。太钢汽水也是尝了尝别的味道，感觉都还可以。</p><p>吃完饭，就去汾河边散散步。这次见识到了传说中的自行车道，其实挺窄的，一个方向只能容纳一个人骑。并且这个自行车道的下匝道有点少，很多地方是高架。要散步的话，可以穿过自行车道再往河边走。</p><p>逛完回去已经快九点半了，赶快打车去晋商博物馆。晋商博物馆实际上就是原来的巡抚衙门，在民国时期也是山西的中枢所在，到了新中国，一度是省委和省政府的办公地，一直到 2017 年改建为博物馆。这个博物馆确实透露着政府办公室的气息，特别是中间一栋苏联式样的办公楼，一进去就能闻到浓烈的复写纸的味道。</p><p>我发现山西这边的博物馆有一个喜好，就是搞大全套收集。例如山西博物院就搞了个钱币展，非常夸张的收集了先秦时期各个国家的钱币，以及后面历朝历代不同年号的钱币。晋商博物馆中更是收集了七八套大小不一的编钟，收集了鼎、盆、壶、斛等各种器皿。最特别的是，我第一次看到了一个东西叫灶。晋商博物馆挺长的，进去先是一个高楼，里面存了据说是山西第一巡抚诺敏的匾额。后面是2号楼和3号楼，两栋楼通过走廊连在一起。在后面是个会议室，和一个苏联样式的高楼。再往后有一个钟楼，但是不让过去了。比较幸运的是，东花园原来在修缮，但是这一次对我们开放了。进去可以看到民国时候的一些陈列，以及五几年的时候省委书记的住房。</p><p>从晋商博物院出来，就去吃了一诺铜火锅。我老婆照例点了一辈子都吃不完的菜，点评如下：</p><ul><li>铜锅 感觉一般</li><li>黄米凉糕 很好吃，但是不要把上面的糖拌进去，不然就太甜了</li><li>丸子串 一般</li><li>羊肉串 很好吃</li><li>涮肚 感觉很好吃，我老婆觉得麻将很多，但是我觉得跟颐祥阁一模一样</li><li>蒜泥茄条 挺好吃的，酸酸甜甜</li><li>小酥肉 肉感觉一般，但是旁边的蘑菇还不错</li><li>烤饼 感觉很好吃，也带回去了</li><li>鸡包豆腐 感觉就是千页豆腐？一般</li><li>过油肉 很好吃</li><li>酸辣白菜 酸酸甜甜的很好吃</li></ul><p>吃完饭，才 12.40，感觉还有一段时间，因此就准备去纯阳宫。因为我在打车过来的路上就看到纯阳宫了，所以就直接骑车过去，风风火火，花了不到十分钟就骑到了。进去之后，我老婆终于要上厕所了，我先进去逛。它的展厅都好没意思，啥陈列都没有，就干放视频。走到最后，是个主题展，里面陈列了一个常阳天尊像，居然是一级国宝，并且是 195 文物。这玩意也没有被框起来，能够近距离观赏感觉挺好，大家也都很有素质，不上手摸。出门，上楼梯，二层的展览就很蠢了，有个展厅里面全是文物已借出。二楼可以通到前面的九宫八卦院，有很多人在那边拍照。</p><p>往回走，打开了旅途随身听，介绍了九宫八卦院，这个建筑布局很有意思，据说是全国唯一的。再往回，就是吕洞宾殿，也就是纯阳宫的主要建筑。再往回就是弥勒佛铜像。我老婆跟我说这里面有个西汉的石狮子，还有明代的铜狮子，居然就直接摆在那里，也没有保护，感觉山西还是富，人们和文物生活在一起。</p><p>再往回走，就快到了出口。此时左边是一个假山，上面是一个关羽像。关羽像没有胡子，因为这个像是在明朝的，而关羽有胡子的说法是明末根据三国演义才有的。右边是个碑廊，然后里面放着另一个 195 文物，涅槃变相碑。这玩意也没有被保护起来，我觉得还是不太好，毕竟这个是室外嘛，还是要有点防护为好。</p><p>我在看涅槃变相碑的时候，因为空间狭小，我的书包还把后面的一个壁画的说明牌给蹭下来了，幸亏没有伤着壁画，不过这也说明了这个陈列比较拥挤，容易碰到。</p><p>从纯阳宫出来立马打车赶往机场，不得再次羡慕太原机场离市区实在是近。</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;总的来说，太原是类似于西安的城市，类似的气候，类似的历史文化。但在游玩体验来讲，太原总体略优于西安，主要是：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;西安商业化太浓重了，例如大唐不夜城、城墙上骑车等。当然这也不是坏事，但如果玩的多了，就会觉得商业化严重的景点如同预制菜一样，不能说不好吃，但感觉容易腻&lt;/li&gt;
&lt;li&gt;西安人太多了&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;但太原的问题主要是：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;交通很不方便。虽然也看到它有专门的旅游线路，并且感觉是用心的，但体验上确实还有欠缺。&lt;/li&gt;
&lt;/ul&gt;</summary>
    
    
    
    
    <category term="游记" scheme="http://www.calvinneo.com/tags/游记/"/>
    
  </entry>
  
  <entry>
    <title>TiCI 上线过程</title>
    <link href="http://www.calvinneo.com/2025/11/30/tici-go-online/"/>
    <id>http://www.calvinneo.com/2025/11/30/tici-go-online/</id>
    <published>2025-11-30T03:57:20.000Z</published>
    <updated>2026-05-12T10:36:43.901Z</updated>
    
    <content type="html"><![CDATA[<p>记录了 TiCI 上线过程中遇到的一些问题：</p><ul><li>针对这些问题的技术性解法和运维性的解法<br>  涉及到某些内部知识的将不予公开。</li><li>对于问题严重程度应该如何判断</li><li>如何为了达成上线的既定目标，设计临时性的缓解措施</li></ul><a id="more"></a><h1 id="POC-on-TiDB-X-for-Customer-U-late-Nov-2025"><a href="#POC-on-TiDB-X-for-Customer-U-late-Nov-2025" class="headerlink" title="POC on TiDB-X for Customer U, late Nov 2025"></a>POC on TiDB-X for Customer U, late Nov 2025</h1><h2 id="主要问题"><a href="#主要问题" class="headerlink" title="主要问题"></a>主要问题</h2><ol><li>支持 Keyspace</li><li>运维手段<br> 包含 Shard、Reader 和 Importer</li><li>支持 Security</li><li>openssl 的编译问题</li><li>arm 编译的问题</li><li>上线后出现的影响可用性的 bug</li><li>上线后出现的影响性能的 bug</li></ol><h2 id="Nov-21"><a href="#Nov-21" class="headerlink" title="Nov 21"></a>Nov 21</h2><p>讨论 import into 场景下，如果某个 worker 因为 OOM 重启，则因为目前缺少 Heartbeat 机制和 reschedule 机制，整个任务会因为这个小任务失败而停止。</p><p>另外，也提到了和 import into 流控相关的问题。</p><h2 id="Nov-23"><a href="#Nov-23" class="headerlink" title="Nov 23"></a>Nov 23</h2><p>遇到了 tokio Sender 报错 channel closed 问题。当时没有空查。</p><p>这个问题实际上就是协程因为哪里 unwrap 而 panic 了，而因为协程池是跑的消息循环，所以 panic 了不 join 也不知道。一直没找到是哪里，后来我加了点 panic hook，才在 Nov 30 定位到是服务发现的问题。</p><h2 id="Nov-26"><a href="#Nov-26" class="headerlink" title="Nov 26"></a>Nov 26</h2><ol><li>TLS options are only supported with HTTPS URLs<br> 这个错误很 plain</li><li>Client had asked for TLS connection but TLS support is disabled. Please enable one of the following features: [“native-tls-tls”, “rusttls-tls”]<br> 我们不能用 native，只能用内置的 rusttls</li><li>undefined symbol: pthread_atfork<br> ARM 上的编译问题，有专门<a href="/2025/11/25/pthread_atfork/">文章</a>介绍</li><li>TLS error error:0A000086<br> 这个就是要启动的时候指定下用 ring 还是 aws_lc_rs</li></ol><h2 id="Nov-27"><a href="#Nov-27" class="headerlink" title="Nov 27"></a>Nov 27</h2><ol><li>排查 ARM 上的问题，这里做了两套方案，先设法用 x86 跑起来。这证明也是正确的，白天我们发现了更多的问题，最终因为 TLS 的问题也没有跑起来。但是晚上我们把 ARM 验证了下，发现不报原先的错了。</li><li>遇到一个新问题是我们的云上 operator 机制不支持改参数，我们挂的又是 ro 文件系统，导致很难验证</li><li>etcd client unavailable or unhealthy, attempting reconnect<br> 这个是 <code>kebab-case</code> 的锅，Rust 和 C++ 的格式不一样，写配置的同学把 ca-path 写成 ca_path 了。</li><li>后来发现，TiFlash CN 还是绕不开 TLS 的问题。所以还是得兼容。</li></ol><h2 id="Nov-28"><a href="#Nov-28" class="headerlink" title="Nov 28"></a>Nov 28</h2><ol><li>get meta channel failed times<br> 这个错误还是 etcd client 的 ca-path 没传对</li><li>到这里为止，整体看上去能跑了，但是查询报错，看日志感觉核心服务没起来<br> 首先 lsof 看了一下，发现服务器都没起来。<br> 因为没有 panic，直觉是哪里有个报错死在协程里面没传播出去。进而发现是 rust 的 grpc server 不能在 dns 格式的 url 上启动。不得不说我几天前的 advertise_addr 的改动很有先见之明 <a href="https://github.com/pingcap-inc/tici/pull/515" target="_blank" rel="noopener">https://github.com/pingcap-inc/tici/pull/515</a></li></ol><h2 id="Nov-29"><a href="#Nov-29" class="headerlink" title="Nov 29"></a>Nov 29</h2><ol><li>白天是在 import into 导入数据，晚上出现了一堆问题。</li><li>首先，是之前的那个 panic 导致协程退出的阴魂不散又来了。我紧急去掉了一堆 unwrap，并且加了 panic hook 打印错误日志。</li><li>然后，是 writer 那边没有对 meta 返回的错误码进行处理，但是加上了处理之后，一个集成测试不通过了。大家觉得要不就把这个测试禁用了吧，但是我很反对，因为这个测试是唯一一个带上真实的 worker 和 reader 跑完 e2e 的测试。并且我看了日志发现一个 writer node 被超时移出了，所以尽管 pr owner 认为他并没有改动到这一块的逻辑，我仍然认为他的修改导致一个严重问题暴露了。因此，我们应该先查问题。原因是即使我们强行合并了，也不能拿这个 commit 去跑生产。后来，确实发现了是在处理 heartbeat 的时候会等待一个异步任务结束，所以导致后面 heartbeat 消息循环直接卡死了。所以，这个问题确实会导致 writer 因为丢失心跳从而被全不 failover，进而整个集群没有 writer 可用的 critical 问题。不 approve 这个 PR 是完全正确的。</li></ol><h2 id="Nov-30"><a href="#Nov-30" class="headerlink" title="Nov 30"></a>Nov 30</h2><p>今天主要切换到性能方面的支持上：</p><ul><li>warmup 的速度太慢了，所以想把这一块改成并发执行<br>  在这个处理之后，8 个节点，24 个并发，大概是不到两个小时就处理完毕了。</li><li>因为 worker 出现了异常重启，导致重启后 meta 向它发送了大量的 add shard message，导致触发了 grpc 的消息大小上限。临时通过增加上限来 workaround 了。实际上是可以通过拆成多个 heartbeat response 来解决。另外同事也提到，可以只发部分信息，其他的后续让主动请求，但这个可能会产生很多的 grpc 调用。</li></ul><h2 id="Dec-1"><a href="#Dec-1" class="headerlink" title="Dec 1"></a>Dec 1</h2><ul><li>发现查询速度比较低，原因是串行访问的所有 shard</li></ul><h2 id="Dec-2"><a href="#Dec-2" class="headerlink" title="Dec 2"></a>Dec 2</h2><ul><li>发现查询并发比较低，原因是查询的 filter 不能通过主键进行过滤，所以导致每次查询需要访问所有的 shard</li></ul><h2 id="Dec-3"><a href="#Dec-3" class="headerlink" title="Dec 3"></a>Dec 3</h2><ul><li>关于昨天的并发问题，认为一个 shard 中有多个 fragment，并且一个 fragment 中又有多个 segment，因此会影响查询的速度。通过 Manual compaction 和调大内存的方式，使得 fragment 和 segment 都变少。</li><li>因为目前没有自动 balance 机制，所以还需要 Manual reschedule。</li></ul><h2 id="Dec-5"><a href="#Dec-5" class="headerlink" title="Dec 5"></a>Dec 5</h2><p>功能侧：</p><ul><li>主要处理 duplicate fragment 的问题。我认为因为 apply compaction 需要根据 frag path 来定位，重名的 frag 会导致问题，所以应该由 meta 来禁用。</li></ul><p>测试侧：</p><ul><li>执行了 Compaction，并且后续自动执行了 Merge，发现查询 QPS 提高了 50%，但依然是比较差的</li><li>因此决定用 (token0_address, ts) 作为新的主键，和 TiKV 解耦</li></ul><h2 id="Dec-6"><a href="#Dec-6" class="headerlink" title="Dec 6"></a>Dec 6</h2><p>这一轮测试发现调整了新的主键之后，QPS 增加了很多，可以看出列存中使用不同的 Sharding key 的重要性：</p><ul><li><p>静态数据 QPS 2.6k+，P999 61.6 ms，且仍有弹性。其中 TiFlash CPU 是 2300%，内存是 8 * 91G。<br>  <a href="https://github.com/CalvinNeo/ue-bench/blob/master/src/main.rs" target="_blank" rel="noopener">https://github.com/CalvinNeo/ue-bench/blob/master/src/main.rs</a></p>  <figure class="highlight sql"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">SELECT</span> * <span class="keyword">FROM</span> ... <span class="keyword">WHERE</span> token0_address = ? <span class="keyword">AND</span> platform = ? <span class="keyword">AND</span> ts &lt; ? <span class="keyword">LIMIT</span> <span class="number">5</span></span><br></pre></td></tr></table></figure></li><li><p>写入的索引大概是 10.5TB 左右。Shard 数量约为 5800+。</p></li></ul><h1 id="某-OP-用户：Q1-2026"><a href="#某-OP-用户：Q1-2026" class="headerlink" title="某 OP 用户：Q1 2026"></a>某 OP 用户：Q1 2026</h1>]]></content>
    
    
    <summary type="html">&lt;p&gt;记录了 TiCI 上线过程中遇到的一些问题：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;针对这些问题的技术性解法和运维性的解法&lt;br&gt;  涉及到某些内部知识的将不予公开。&lt;/li&gt;
&lt;li&gt;对于问题严重程度应该如何判断&lt;/li&gt;
&lt;li&gt;如何为了达成上线的既定目标，设计临时性的缓解措施&lt;/li&gt;
&lt;/ul&gt;</summary>
    
    
    
    
    <category term="数据库" scheme="http://www.calvinneo.com/tags/数据库/"/>
    
    <category term="Rust" scheme="http://www.calvinneo.com/tags/Rust/"/>
    
  </entry>
  
  <entry>
    <title>undefined symbol pthread_atfork</title>
    <link href="http://www.calvinneo.com/2025/11/25/pthread_atfork/"/>
    <id>http://www.calvinneo.com/2025/11/25/pthread_atfork/</id>
    <published>2025-11-24T18:20:13.000Z</published>
    <updated>2025-12-03T08:21:21.643Z</updated>
    
    <content type="html"><![CDATA[<p>在 x86 上可以跑，但是在 arm linux 上就报这个错误。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">/tiflash/tiflash: symbol lookup error: /tiflash/libtici_search_lib.so: undefined symbol: pthread_atfork</span><br></pre></td></tr></table></figure><a id="more"></a><p>首先，ldd 看到，链接的是本地的 <code>/lib64/libpthread.so.0</code>。</p><p>可以通过 <code>strings /lib64/libpthread.so.0 | grep &#39;^GLIBC_&#39;</code> 命令查询 GLIBC 的版本。</p><p>然后，nm 了一下 /tiflash/libtici_search_lib.so，结果是：</p><ol><li><p>x86 开发机</p> <figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br></pre></td><td class="code"><pre><span class="line">nm libtici_search_lib.so | grep pthread</span><br><span class="line">0000000003c97e70 t __pthread_atfork</span><br><span class="line">0000000003c97e70 t pthread_atfork</span><br><span class="line">                 U pthread_attr_destroy</span><br><span class="line">                 U pthread_attr_getguardsize</span><br><span class="line">                 U pthread_attr_getstack</span><br><span class="line">                 U pthread_attr_init</span><br><span class="line">                 U pthread_attr_setstacksize</span><br></pre></td></tr></table></figure></li><li><p>arm tiflash:v2025.8.10-2-gc9e3144-centos7 镜像<br> <img src="/img/pthread_atfork/arm.jpg"></p></li><li><p>x86 tiflash:v2025.8.10-2-gc9e3144-centos7 镜像<br> 这个是 multi arch 镜像，但是 glibc 版本不一样</p></li></ol><p>那么这个符号在 <code>/lib64/libpthread.so.0</code> 里面有么？nm 了一下：</p><ul><li>x86 的版本是 2.34，显示这个 so 是没有 Debug info 的。难道被 trim 了么？ls 了一下这个文件，发现只有 15KiB 左右。后来了解到，在较新的 GLIBC 中，pthread 相关的被整合到了 libc.so 中，我 nm 了 libc.so 确实可以看到。</li><li>arm 版本是 2.17，nm 了可以看到其他 pthread 符号，但是看不到 pthread_atfork。</li></ul><p><img src="/img/pthread_atfork/arm33.png"></p><p>原因是在大多数 Linux 发行版（使用 glibc 的系统）中，pthread_atfork 的符号并不在常规的共享库如 libpthread.so.0 中，而是通过链接器脚本特殊处理，其具体实现位于 libpthread_nonshared.a 这个静态归档文件中，而非 .so 文件里 。</p><p>解决方案参考 <a href="https://github.com/pingcap/tiflash/pull/10571%EF%BC%8C%E5%BC%BA%E5%88%B6%E4%BD%BF%E7%94%A8" target="_blank" rel="noopener">https://github.com/pingcap/tiflash/pull/10571，强制使用</a> <code>-pthread</code> 而不是 <code>-lpthread</code> 即可。</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;在 x86 上可以跑，但是在 arm linux 上就报这个错误。&lt;/p&gt;
&lt;figure class=&quot;highlight plain&quot;&gt;&lt;table&gt;&lt;tr&gt;&lt;td class=&quot;gutter&quot;&gt;&lt;pre&gt;&lt;span class=&quot;line&quot;&gt;1&lt;/span&gt;&lt;br&gt;&lt;/pre&gt;&lt;/td&gt;&lt;td class=&quot;code&quot;&gt;&lt;pre&gt;&lt;span class=&quot;line&quot;&gt;/tiflash/tiflash: symbol lookup error: /tiflash/libtici_search_lib.so: undefined symbol: pthread_atfork&lt;/span&gt;&lt;br&gt;&lt;/pre&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;&lt;/figure&gt;</summary>
    
    
    
    
    <category term="C++" scheme="http://www.calvinneo.com/tags/C/"/>
    
    <category term="Rust" scheme="http://www.calvinneo.com/tags/Rust/"/>
    
    <category term="GLIBC" scheme="http://www.calvinneo.com/tags/GLIBC/"/>
    
  </entry>
  
  <entry>
    <title>tokio channel 实现</title>
    <link href="http://www.calvinneo.com/2025/11/23/tokio_channel_src/"/>
    <id>http://www.calvinneo.com/2025/11/23/tokio_channel_src/</id>
    <published>2025-11-23T15:09:06.000Z</published>
    <updated>2025-12-12T14:50:31.126Z</updated>
    
    <content type="html"><![CDATA[<p>基于 tokio 1.46.0 版本</p><a id="more"></a><h1 id="mpsc"><a href="#mpsc" class="headerlink" title="mpsc"></a>mpsc</h1><p>mpsc 有 bounded 和 unbounded 两种形式。通过不同的 semaphore 来区别。</p><h2 id="Chan-结构"><a href="#Chan-结构" class="headerlink" title="Chan 结构"></a>Chan 结构</h2><p>对于 unbounded semaphore，其最低的 bit 表示这个 channel 有没有关闭。</p><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br><span class="line">20</span><br><span class="line">21</span><br><span class="line">22</span><br><span class="line">23</span><br><span class="line">24</span><br><span class="line">25</span><br><span class="line">26</span><br><span class="line">27</span><br><span class="line">28</span><br><span class="line">29</span><br><span class="line">30</span><br><span class="line">31</span><br><span class="line">32</span><br><span class="line">33</span><br><span class="line">34</span><br><span class="line">35</span><br><span class="line">36</span><br><span class="line">37</span><br><span class="line">38</span><br><span class="line">39</span><br><span class="line">40</span><br><span class="line">41</span><br><span class="line">42</span><br><span class="line">43</span><br><span class="line">44</span><br><span class="line">45</span><br><span class="line">46</span><br><span class="line">47</span><br><span class="line">48</span><br><span class="line">49</span><br><span class="line">50</span><br><span class="line">51</span><br><span class="line">52</span><br><span class="line">53</span><br><span class="line">54</span><br><span class="line">55</span><br><span class="line">56</span><br><span class="line">57</span><br></pre></td><td class="code"><pre><span class="line"><span class="comment">// ===== impl Semaphore for (::Semaphore, capacity) =====</span></span><br><span class="line"></span><br><span class="line"><span class="keyword">impl</span> Semaphore <span class="keyword">for</span> bounded::Semaphore &#123;</span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">add_permit</span></span>(&amp;<span class="keyword">self</span>) &#123;</span><br><span class="line">        <span class="keyword">self</span>.semaphore.release(<span class="number">1</span>);</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">add_permits</span></span>(&amp;<span class="keyword">self</span>, n: <span class="built_in">usize</span>) &#123;</span><br><span class="line">        <span class="keyword">self</span>.semaphore.release(n)</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">is_idle</span></span>(&amp;<span class="keyword">self</span>) -&gt; <span class="built_in">bool</span> &#123;</span><br><span class="line">        <span class="keyword">self</span>.semaphore.available_permits() == <span class="keyword">self</span>.bound</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">close</span></span>(&amp;<span class="keyword">self</span>) &#123;</span><br><span class="line">        <span class="keyword">self</span>.semaphore.close();</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">is_closed</span></span>(&amp;<span class="keyword">self</span>) -&gt; <span class="built_in">bool</span> &#123;</span><br><span class="line">        <span class="keyword">self</span>.semaphore.is_closed()</span><br><span class="line">    &#125;</span><br><span class="line">&#125;</span><br><span class="line"></span><br><span class="line"><span class="comment">// ===== impl Semaphore for AtomicUsize =====</span></span><br><span class="line"></span><br><span class="line"><span class="keyword">impl</span> Semaphore <span class="keyword">for</span> unbounded::Semaphore &#123;</span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">add_permit</span></span>(&amp;<span class="keyword">self</span>) &#123;</span><br><span class="line">        <span class="keyword">let</span> prev = <span class="keyword">self</span>.<span class="number">0</span>.fetch_sub(<span class="number">2</span>, Release);</span><br><span class="line"></span><br><span class="line">        <span class="keyword">if</span> prev &gt;&gt; <span class="number">1</span> == <span class="number">0</span> &#123;</span><br><span class="line">            <span class="comment">// Something went wrong</span></span><br><span class="line">            process::abort();</span><br><span class="line">        &#125;</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">add_permits</span></span>(&amp;<span class="keyword">self</span>, n: <span class="built_in">usize</span>) &#123;</span><br><span class="line">        <span class="keyword">let</span> prev = <span class="keyword">self</span>.<span class="number">0</span>.fetch_sub(n &lt;&lt; <span class="number">1</span>, Release);</span><br><span class="line"></span><br><span class="line">        <span class="keyword">if</span> (prev &gt;&gt; <span class="number">1</span>) &lt; n &#123;</span><br><span class="line">            <span class="comment">// Something went wrong</span></span><br><span class="line">            process::abort();</span><br><span class="line">        &#125;</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">is_idle</span></span>(&amp;<span class="keyword">self</span>) -&gt; <span class="built_in">bool</span> &#123;</span><br><span class="line">        <span class="keyword">self</span>.<span class="number">0</span>.load(Acquire) &gt;&gt; <span class="number">1</span> == <span class="number">0</span></span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">close</span></span>(&amp;<span class="keyword">self</span>) &#123;</span><br><span class="line">        <span class="keyword">self</span>.<span class="number">0</span>.fetch_or(<span class="number">1</span>, Release);</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="function"><span class="keyword">fn</span> <span class="title">is_closed</span></span>(&amp;<span class="keyword">self</span>) -&gt; <span class="built_in">bool</span> &#123;</span><br><span class="line">        <span class="keyword">self</span>.<span class="number">0</span>.load(Acquire) &amp; <span class="number">1</span> == <span class="number">1</span></span><br><span class="line">    &#125;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>无论是 unbounded 还是 bounded，最终都是到 chan::Tx 和 chan::Rx 这两个结构里面。这两个类都持有一个 <code>Arc&lt;Chan&lt;T, S&gt;&gt;</code>。</p><p>我们会看到有两个 Tx：</p><ol><li>chan::Tx 是一个 <code>Arc&lt;Chan&lt;T, S&gt;&gt;</code>，它实际上封装了 Chan，更上层一点</li><li>list::Tx 是一个 Block 的链表。它是 <code>Chan&lt;T, S&gt;</code> 的 filed，更底层一点</li></ol><p>list::tx 是一个无锁队列。这个队列的内存是以 Block 为基础分配的，每个 block 能装 <code>const BLOCK_CAP: usize = 32;</code> 个 Value。所以，Chan 中实现了解耦：</p><ul><li>list 模块只负责无锁队列的实现</li><li>Chan 的其他部分负责容量控制、notify/waker 的逻辑</li></ul><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br><span class="line">20</span><br><span class="line">21</span><br><span class="line">22</span><br><span class="line">23</span><br><span class="line">24</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">pub</span>(<span class="keyword">super</span>) <span class="class"><span class="keyword">struct</span> <span class="title">Chan</span></span>&lt;T, S&gt; &#123;</span><br><span class="line">    <span class="comment">/// Handle to the push half of the lock-free list.</span></span><br><span class="line">    tx: CachePadded&lt;list::Tx&lt;T&gt;&gt;,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// Receiver waker. Notified when a value is pushed into the channel.</span></span><br><span class="line">    rx_waker: CachePadded&lt;AtomicWaker&gt;,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// Notifies all tasks listening for the receiver being dropped.</span></span><br><span class="line">    notify_rx_closed: Notify,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// Coordinates access to channel's capacity.</span></span><br><span class="line">    semaphore: S,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// Tracks the number of outstanding sender handles.</span></span><br><span class="line">    <span class="comment">///</span></span><br><span class="line">    <span class="comment">/// When this drops to zero, the send half of the channel is closed.</span></span><br><span class="line">    tx_count: AtomicUsize,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// Tracks the number of outstanding weak sender handles.</span></span><br><span class="line">    tx_weak_count: AtomicUsize,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// Only accessed by `Rx` handle.</span></span><br><span class="line">    rx_fields: UnsafeCell&lt;RxFields&lt;T&gt;&gt;,</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><h3 id="block-的实现"><a href="#block-的实现" class="headerlink" title="block 的实现"></a>block 的实现</h3><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br><span class="line">20</span><br><span class="line">21</span><br><span class="line">22</span><br><span class="line">23</span><br><span class="line">24</span><br><span class="line">25</span><br><span class="line">26</span><br><span class="line">27</span><br><span class="line">28</span><br><span class="line">29</span><br><span class="line">30</span><br><span class="line">31</span><br><span class="line">32</span><br><span class="line">33</span><br></pre></td><td class="code"><pre><span class="line"><span class="meta">#[cfg(all(target_pointer_width = <span class="meta-string">"64"</span>, not(loom)))]</span></span><br><span class="line"><span class="keyword">const</span> BLOCK_CAP: <span class="built_in">usize</span> = <span class="number">32</span>;</span><br><span class="line"></span><br><span class="line"><span class="meta">#[repr(transparent)]</span></span><br><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">Values</span></span>&lt;T&gt;([UnsafeCell&lt;MaybeUninit&lt;T&gt;&gt;; BLOCK_CAP]);</span><br><span class="line"></span><br><span class="line"><span class="keyword">pub</span>(<span class="keyword">crate</span>) <span class="class"><span class="keyword">struct</span> <span class="title">Block</span></span>&lt;T&gt; &#123;</span><br><span class="line">    <span class="comment">/// The header fields.</span></span><br><span class="line">    header: BlockHeader&lt;T&gt;,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// Array containing values pushed into the block. Values are stored in a</span></span><br><span class="line">    <span class="comment">/// continuous array in order to improve cache line behavior when reading.</span></span><br><span class="line">    <span class="comment">/// The values must be manually dropped.</span></span><br><span class="line">    values: Values&lt;T&gt;,</span><br><span class="line">&#125;</span><br><span class="line"></span><br><span class="line"><span class="comment">/// Extra fields for a `Block&lt;T&gt;`.</span></span><br><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">BlockHeader</span></span>&lt;T&gt; &#123;</span><br><span class="line">    <span class="comment">/// The start index of this block.</span></span><br><span class="line">    <span class="comment">///</span></span><br><span class="line">    <span class="comment">/// Slots in this block have indices in `start_index .. start_index + BLOCK_CAP`.</span></span><br><span class="line">    start_index: <span class="built_in">usize</span>,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// The next block in the linked list.</span></span><br><span class="line">    next: AtomicPtr&lt;Block&lt;T&gt;&gt;,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// Bitfield tracking slots that are ready to have their values consumed.</span></span><br><span class="line">    ready_slots: AtomicUsize,</span><br><span class="line"></span><br><span class="line">    <span class="comment">/// The observed `tail_position` value *after* the block has been passed by</span></span><br><span class="line">    <span class="comment">/// `block_tail`.</span></span><br><span class="line">    observed_tail_position: UnsafeCell&lt;<span class="built_in">usize</span>&gt;,</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>看 Block::grow，就是一个无锁链表的实现 </p><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br></pre></td><td class="code"><pre><span class="line">...</span><br><span class="line">        <span class="keyword">let</span> <span class="keyword">mut</span> curr = next;</span><br><span class="line"></span><br><span class="line">        <span class="comment">// <span class="doctag">TODO:</span> Should this iteration be capped?</span></span><br><span class="line">        <span class="keyword">loop</span> &#123;</span><br><span class="line">            <span class="keyword">let</span> actual = <span class="keyword">unsafe</span> &#123; curr.as_ref().try_push(&amp;<span class="keyword">mut</span> new_block, AcqRel, Acquire) &#125;;</span><br><span class="line"></span><br><span class="line">            curr = <span class="keyword">match</span> actual &#123;</span><br><span class="line">                <span class="literal">Ok</span>(()) =&gt; &#123;</span><br><span class="line">                    <span class="keyword">return</span> next;</span><br><span class="line">                &#125;</span><br><span class="line">                <span class="literal">Err</span>(curr) =&gt; curr,</span><br><span class="line">            &#125;;</span><br><span class="line"></span><br><span class="line">            crate::loom::thread::yield_now();</span><br><span class="line">        &#125;</span><br><span class="line">...</span><br></pre></td></tr></table></figure><h4 id="为什么是关于-Block-的无锁队列？"><a href="#为什么是关于-Block-的无锁队列？" class="headerlink" title="为什么是关于 Block 的无锁队列？"></a>为什么是关于 Block 的无锁队列？</h4><blockquote><p>为什么是关于 Block 的无锁队列，而不是关于 Value 的呢？</p></blockquote><ol><li>减少锁竞争和原子操作 (Reduced Contention and Atomic Operations)<br> 集中操作 (Batch Operations): 在高并发场景下，如果每次发送一个 Value 就需要对共享队列（如链表或原子指针）进行一次原子操作（例如 CAS - Compare-and-Swap）来添加节点，那么多个发送者 (Multi-Producer) 之间的竞争会非常激烈，导致性能瓶颈。<br> Block 机制: 使用 Block（块），每个 Block 中可以存放多个 Value。这样，发送者可以一次性分配一个 Block，并填充多个 Value。在将这个 Block 链接到队列末尾时，只需要进行一次原子操作来更新队列的尾指针。这大大减少了对核心共享数据结构（队列头/尾）的原子操作次数，从而降低了锁竞争和系统开销。<br> 局部性 (Locality): 一旦一个发送者成功获取并填充了一个 Block，它就可以在没有竞争的情况下，向该 Block 中写入若干条消息，这利用了 CPU 缓存的局部性原理。</li><li>优化内存分配 (Optimized Memory Allocation)<br> 批量分配: 每次发送一个 Value 就进行一次内存分配是低效的。Block 允许一次性分配一块较大的内存，用于存储多个 Value。<br> 更少的元数据: 如果每条 Value 都是一个独立的链表节点，那么每个节点都需要存储一个指针（指向下一个节点）作为元数据。在 Block 方案中，只有 Block 之间有链接指针，一个 Block 内部的多个 Value 可以紧凑存储，减少了内存开销。</li></ol><h2 id="channel-创建"><a href="#channel-创建" class="headerlink" title="channel 创建"></a>channel 创建</h2><h2 id="send-实现"><a href="#send-实现" class="headerlink" title="send 实现"></a>send 实现</h2><h3 id="unbounded"><a href="#unbounded" class="headerlink" title="unbounded"></a>unbounded</h3><p>unbounded 的 send 的实现，可以看到，最终调用了 Chan 的 send，后面会详细介绍。</p><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">pub</span> <span class="function"><span class="keyword">fn</span> <span class="title">send</span></span>(&amp;<span class="keyword">self</span>, message: T) -&gt; <span class="built_in">Result</span>&lt;(), SendError&lt;T&gt;&gt; &#123;</span><br><span class="line">    <span class="keyword">if</span> !<span class="keyword">self</span>.inc_num_messages() &#123;</span><br><span class="line">        <span class="keyword">return</span> <span class="literal">Err</span>(SendError(message));</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="keyword">self</span>.chan.send(message);</span><br><span class="line">    <span class="literal">Ok</span>(())</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>inc_num_messages 的逻辑如下，主要是每次增加 2，保证末尾是 0</p><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">if</span> curr &amp; <span class="number">1</span> == <span class="number">1</span> &#123;</span><br><span class="line"><span class="keyword">return</span> <span class="literal">false</span>;</span><br><span class="line">&#125;</span><br><span class="line"></span><br><span class="line">.compare_exchange(curr, curr + <span class="number">2</span>, AcqRel, Acquire)</span><br></pre></td></tr></table></figure><h3 id="bounded"><a href="#bounded" class="headerlink" title="bounded"></a>bounded</h3><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br></pre></td><td class="code"><pre><span class="line">async <span class="function"><span class="keyword">fn</span> <span class="title">reserve_inner</span></span>(&amp;<span class="keyword">self</span>, n: <span class="built_in">usize</span>) -&gt; <span class="built_in">Result</span>&lt;(), SendError&lt;()&gt;&gt; &#123;</span><br><span class="line">    crate::trace::async_trace_leaf().await;</span><br><span class="line"></span><br><span class="line">    <span class="keyword">if</span> n &gt; <span class="keyword">self</span>.max_capacity() &#123;</span><br><span class="line">        <span class="keyword">return</span> <span class="literal">Err</span>(SendError(()));</span><br><span class="line">    &#125;</span><br><span class="line">    <span class="keyword">match</span> <span class="keyword">self</span>.chan.semaphore().semaphore.acquire(n).await &#123;</span><br><span class="line">        <span class="literal">Ok</span>(()) =&gt; <span class="literal">Ok</span>(()),</span><br><span class="line">        <span class="literal">Err</span>(_) =&gt; <span class="literal">Err</span>(SendError(())),</span><br><span class="line">    &#125;</span><br><span class="line">&#125;</span><br><span class="line"></span><br><span class="line"><span class="keyword">pub</span> <span class="function"><span class="keyword">fn</span> <span class="title">max_capacity</span></span>(&amp;<span class="keyword">self</span>) -&gt; <span class="built_in">usize</span> &#123;</span><br><span class="line">    <span class="keyword">self</span>.chan.semaphore().bound</span><br><span class="line">&#125;</span><br><span class="line"></span><br><span class="line"><span class="keyword">pub</span>(<span class="keyword">crate</span>) <span class="function"><span class="keyword">fn</span> <span class="title">acquire</span></span>(&amp;<span class="keyword">self</span>, num_permits: <span class="built_in">usize</span>) -&gt; Acquire&lt;<span class="symbol">'_</span>&gt; &#123;</span><br><span class="line">    Acquire::new(<span class="keyword">self</span>, num_permits)</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><h3 id="Chan-的-send-实现"><a href="#Chan-的-send-实现" class="headerlink" title="Chan 的 send 实现"></a>Chan 的 send 实现</h3><p>Chan 的 send 的实现如下</p><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br></pre></td><td class="code"><pre><span class="line"><span class="function"><span class="keyword">fn</span> <span class="title">send</span></span>(&amp;<span class="keyword">self</span>, value: T) &#123;</span><br><span class="line">    <span class="comment">// Push the value</span></span><br><span class="line">    <span class="keyword">self</span>.tx.push(value);</span><br><span class="line"></span><br><span class="line">    <span class="comment">// Notify the rx task</span></span><br><span class="line">    <span class="keyword">self</span>.rx_waker.wake();</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>这个 rx_waker 是一个 AtomicWaker 对象。</p><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">pub</span>(<span class="keyword">crate</span>) <span class="class"><span class="keyword">struct</span> <span class="title">AtomicWaker</span></span> &#123;</span><br><span class="line">    state: AtomicUsize,</span><br><span class="line">    waker: UnsafeCell&lt;<span class="built_in">Option</span>&lt;std::task::Waker&gt;&gt;,</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><h1 id="Reference"><a href="#Reference" class="headerlink" title="Reference"></a>Reference</h1>]]></content>
    
    
    <summary type="html">&lt;p&gt;基于 tokio 1.46.0 版本&lt;/p&gt;</summary>
    
    
    
    
    <category term="Rust" scheme="http://www.calvinneo.com/tags/Rust/"/>
    
  </entry>
  
  <entry>
    <title>“良定义”的状态</title>
    <link href="http://www.calvinneo.com/2025/11/20/well-defined-state/"/>
    <id>http://www.calvinneo.com/2025/11/20/well-defined-state/</id>
    <published>2025-11-20T15:09:06.000Z</published>
    <updated>2025-11-21T15:52:10.678Z</updated>
    
    <content type="html"><![CDATA[<p>我觉得“良定义”的状态需要具备：</p><ul><li>完备性</li><li>唯一性</li></ul><a id="more"></a><h1 id="Case-by-case"><a href="#Case-by-case" class="headerlink" title="Case by case"></a>Case by case</h1><h2 id="如何表示一个空区间？"><a href="#如何表示一个空区间？" class="headerlink" title="如何表示一个空区间？"></a>如何表示一个空区间？</h2><p>有一些 string，那么可以用 [s1, s2) 来圈出一些 string。因为是左闭右开区间，所以空集难以表示。这里只能通过</p><h2 id="如何表示“和上次一样”"><a href="#如何表示“和上次一样”" class="headerlink" title="如何表示“和上次一样”"></a>如何表示“和上次一样”</h2><h2 id="Option-lt-Vec-gt-还是-Vec？"><a href="#Option-lt-Vec-gt-还是-Vec？" class="headerlink" title="Option&lt;Vec&gt; 还是 Vec？"></a>Option&lt;Vec<t>&gt; 还是 Vec<t>？</t></t></h2><p>Vec 是 Option 的上位类型。</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;我觉得“良定义”的状态需要具备：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;完备性&lt;/li&gt;
&lt;li&gt;唯一性&lt;/li&gt;
&lt;/ul&gt;</summary>
    
    
    
    
    <category term="数据结构" scheme="http://www.calvinneo.com/tags/数据结构/"/>
    
    <category term="编程思想" scheme="http://www.calvinneo.com/tags/编程思想/"/>
    
  </entry>
  
  <entry>
    <title>Efficient IO with io_uring 学习</title>
    <link href="http://www.calvinneo.com/2025/10/30/efficient-io-uring/"/>
    <id>http://www.calvinneo.com/2025/10/30/efficient-io-uring/</id>
    <published>2025-10-30T14:20:13.000Z</published>
    <updated>2025-11-06T15:16:59.866Z</updated>
    
    <content type="html"><![CDATA[<p>通过 <a href="https://kernel.dk/io_uring.pdf" target="_blank" rel="noopener">https://kernel.dk/io_uring.pdf</a> 简单学习下 io_uring。</p><a id="more"></a><h1 id="1-0-Introduction"><a href="#1-0-Introduction" class="headerlink" title="1.0 Introduction"></a>1.0 Introduction</h1><p>Linux 的读写 API 经历了：</p><ul><li><p>read</p></li><li><p>pread：增加了 offset</p></li><li><p>preadv：offset 是 iovec 的形式，就是支持分散读</p>  <figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">iovec</span></span></span><br><span class="line"><span class="class">&#123;</span></span><br><span class="line">    <span class="keyword">void</span> __user *iov_base;</span><br><span class="line">    <span class="keyword">__kernel_size_t</span> iov_len;</span><br><span class="line">&#125;;</span><br></pre></td></tr></table></figure></li><li><p>preadv2：加了 flags<br>  可以参考 <a href="https://www.man7.org/linux/man-pages/man2/preadv2.2.html" target="_blank" rel="noopener">https://www.man7.org/linux/man-pages/man2/preadv2.2.html</a></p></li></ul><p>但上述的 API 都是同步的。posix 有 <code>aio_</code> 系列的 API 标准，但是没啥人用，性能也不好。</p><p>Linux 有个 libaio，它和 POSIX 的 <code>aio_</code> 系列不是一个东西。但它也有问题：</p><ul><li><p>它要求 O_DIRECT，不然就和同步调用没啥区别。而 O_DIRECT 会 bypass cache，并且有严格的对齐要求，所以用途受限制。</p></li><li><p>即使满足 async 的所有条件，最终也不一定是 async 的。比如：</p><ul><li>如果要修改元数据，可能会 block</li><li>storage device 的 request slots 的数量是固定的<br>  这里的 request slots 表示 storage device 同时可以处理的并发数。<br>  传统存储协议如 SATA、SAS 中，只有一个命令队列，存放未完成的 io，它的长度就是 io depth。如果下层 storage device 的 request slots 数量小于 io depth，那么 io 请求就可能在 io 队列中等待。<br>  NVMe SSD 支持多个 Submission Queues (SQ) 和 Completion Queues (CQ)，每个 SQ 条目可对应一个正在执行的 I/O 命令。比如有 64 个 queue，每个 queue 深度是 1024，那么理论上最多可并行执行 64 × 1024 = 65536 个命令。</li></ul></li><li><p>提交一个 io 需要复制 64 + 8 bytes。完成一个 io 需要复制 32 bytes。</p>  <figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br><span class="line">20</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">iocb</span> &#123;</span></span><br><span class="line">   __u64   aio_data;</span><br><span class="line">   __<span class="function">u32   <span class="title">PADDED</span><span class="params">(aio_key, aio_rw_flags)</span></span>;</span><br><span class="line">   __u16   aio_lio_opcode;</span><br><span class="line">   __s16   aio_reqprio;</span><br><span class="line">   __u32   aio_fildes;</span><br><span class="line">   __u64   aio_buf;</span><br><span class="line">   __u64   aio_nbytes;</span><br><span class="line">   __s64   aio_offset;</span><br><span class="line">   __u64   aio_reserved2;</span><br><span class="line">   __u32   aio_flags;</span><br><span class="line">   __u32   aio_resfd;</span><br><span class="line">&#125;;</span><br><span class="line"></span><br><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">io_event</span> &#123;</span></span><br><span class="line">    __u64   data;</span><br><span class="line">    __u64   obj;</span><br><span class="line">    __s64   res;</span><br><span class="line">    __s64   res2;</span><br><span class="line">&#125;;</span><br></pre></td></tr></table></figure><p>  Depending on your IO size, this can definitely be noticeable.<br>  IO always requires at least two system calls (submit + wait-for-completion), which in these post spectre/meltdown days is a serious slowdown.</p></li></ul><h1 id="2-0-Improving-the-status-quo"><a href="#2-0-Improving-the-status-quo" class="headerlink" title="2.0 Improving the status quo"></a>2.0 Improving the status quo</h1><p>一开始有一些改良 libaio 的工作：</p><ul><li>If you can extend and improve an existing interface, that’s preferable to providing a new one.</li><li>It’s a lot less work in general.</li></ul><p>libaio 主要有三个接口：</p><ul><li>io_setup</li><li>io_submit 用来提交一个 io</li><li>io_getevents 用来等待完成，并收获结果</li></ul><p>后面觉得，这种改良会把接口改得非常复杂，而且只能解决上面列出的一个问题。</p><h1 id="3-0-New-interface-design-goals"><a href="#3-0-New-interface-design-goals" class="headerlink" title="3.0 New interface design goals"></a>3.0 New interface design goals</h1><ul><li>Easy to use, hard to misuse.</li><li>Extendable. 希望这个接口不止支持 block oriented IO。对于网络，和非块存储设备，它都能适用。</li><li>Feature rich. Linux aio caters to a subset (of a subset) of applications. I did not want to create yet another<br>interface that only covered some of what applications need, or that required applications to reinvent the same<br>functionality over and over again (like IO thread pools).</li><li>Efficiency. While storage IO is mostly still block based and hence at least 512b or 4kb in size, efficiency at those<br>sizes is still critical for certain applications. Additionally, some requests may not even be carrying a data payload.<br>It was important that the new interface was efficient in terms of per-request overhead.</li><li>Scalability. While efficiency and low latencies are important, it’s also critical to provide the best performance<br>possible at the peak end. For storage in particular, we’ve worked very hard to deliver a scalable infrastructure. A<br>new interface should allow us to expose that scalability all the way back to applications.</li></ul><h1 id="4-0-Enter-io-uring"><a href="#4-0-Enter-io-uring" class="headerlink" title="4.0 Enter io_uring"></a>4.0 Enter io_uring</h1><p>首先是摘录作者的感言，性能必须从一开始，在<strong>设计接口</strong>的时候就考虑。</p><blockquote><p>Despite the ranked list of design goals, the initial design was centered around efficiency. Efficiency isn’t something that can be an afterthought, it has to be designed in from the start - you can’t wring it out of something later on once the interface is fixed.</p></blockquote><p>作者认为，新的设计要避免 submission 和 completion 事件在内核和用户空间之间的复制，也要避免 indirection，所以他由浅及深得出了下面几点：</p><ol><li>内核和用户空间需要 share 这些结构</li><li>因此，这些结构应该在内核和用户的共享内存中</li><li>因此，必须要去维护这里面的同步关系</li><li>如果要用锁，那么就肯定会有系统调用，系统调用肯定 overhead 就大了</li><li>因此，single producer single consumer ring buffer 是适合的</li></ol><p>考虑到对于 submission 事件，用户是生产者，内核是消费者；而 completion 事件则相反。所以需要两个队列：SQ 和 CQ。</p><h2 id="4-1-DATA-STRUCTURES"><a href="#4-1-DATA-STRUCTURES" class="headerlink" title="4.1 DATA STRUCTURES"></a>4.1 DATA STRUCTURES</h2><p>cqe 的后缀表示 Completion Queue Event。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">io_uring_cqe</span> &#123;</span></span><br><span class="line">   <span class="comment">// 从 submission 中透传过来</span></span><br><span class="line">   __u64 user_data;</span><br><span class="line">   __s32 res;</span><br><span class="line">   __u32 flags;</span><br><span class="line">&#125;;</span><br></pre></td></tr></table></figure><p>sqe 则复杂很多</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br><span class="line">20</span><br><span class="line">21</span><br><span class="line">22</span><br><span class="line">23</span><br><span class="line">24</span><br><span class="line">25</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">io_uring_sqe</span> &#123;</span></span><br><span class="line">   <span class="comment">// 操作类型，例如 IORING_OP_READV 表示向量读</span></span><br><span class="line">   __u8 opcode;</span><br><span class="line">   __u8 flags;</span><br><span class="line">   __u16 ioprio;</span><br><span class="line">   __s32 fd;</span><br><span class="line">   __u64 off;</span><br><span class="line">   <span class="comment">// 指向内存地址，如果是向量读写，则指向一个 iovec array 的地址</span></span><br><span class="line">   __u64 addr;</span><br><span class="line">   <span class="comment">// 表示长度，或者 iovec array 的长度</span></span><br><span class="line">   __u32 len;</span><br><span class="line">   <span class="keyword">union</span> &#123;</span><br><span class="line">      <span class="keyword">__kernel_rwf_t</span> rw_flags;</span><br><span class="line">      __u32 fsync_flags;</span><br><span class="line">      __u16 poll_events;</span><br><span class="line">      __u32 sync_range_flags;</span><br><span class="line">      __u32 msg_flags;   </span><br><span class="line">   &#125;;</span><br><span class="line">   __u64 user_data;</span><br><span class="line">   <span class="keyword">union</span> &#123;</span><br><span class="line">      __u16 buf_index;</span><br><span class="line">      <span class="comment">// 64 bytes 对齐</span></span><br><span class="line">      __u64 __pad2[<span class="number">3</span>];</span><br><span class="line">   &#125;;</span><br><span class="line">&#125;;</span><br></pre></td></tr></table></figure><h2 id="4-2-COMMUNICATION-CHANNEL"><a href="#4-2-COMMUNICATION-CHANNEL" class="headerlink" title="4.2 COMMUNICATION CHANNEL"></a>4.2 COMMUNICATION CHANNEL</h2><p>SQ 和 CQ 的 indexing 是不太一样的，先从简单的 CQ 开始。</p><p>cqe 是一个内核和用户共享的 ring buffer，内核写会更新 tail，用户读会更新 head。ring buffer 的大小是 2 的幂，它的好处我在 <a href="/2018/07/23/redis_learn_object/">Redis底层对象实现原理分析</a>中有所解析。</p><p>如下所示，head 是可以自然溢出的。当然，正如我在 <a href="/2017/12/05/libutp%E6%BA%90%E7%A0%81%E7%AE%80%E6%9E%90/">libutp源码简析</a>或者<a href="https://github.com/calvinneo/atp" target="_blank" rel="noopener">ATP</a>中的实现那样，当 tail 比 head 小的时候，我们也可以认为发生了溢出。<br><code>cqring-&gt;cqes</code> 是被共享的结构。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">unsigned</span> head;</span><br><span class="line">head = cqring-&gt;head;</span><br><span class="line">read_barrier();</span><br><span class="line"><span class="keyword">if</span> (head != cqring-&gt;tail) &#123;</span><br><span class="line">   <span class="class"><span class="keyword">struct</span> <span class="title">io_uring_cqe</span> *<span class="title">cqe</span>;</span></span><br><span class="line">   <span class="keyword">unsigned</span> index;</span><br><span class="line">   index = head &amp; (cqring-&gt;mask);</span><br><span class="line">   cqe = &amp;cqring-&gt;cqes[index];</span><br><span class="line">   <span class="comment">/* process completed cqe here */</span></span><br><span class="line">   ...</span><br><span class="line">   <span class="comment">/* we've now consumed this entry */</span></span><br><span class="line">   head++;</span><br><span class="line">&#125;</span><br><span class="line">cqring-&gt;head = head;</span><br><span class="line">write_barrier();</span><br></pre></td></tr></table></figure><p>SQ 这边，就是用户生产，内核消费了。之前说到，SQ 的 indexing 不一样，它是有个 indirection 的。submission 的 ring buffer 中存放了 index，索引到 sqe 中的位置。例如下面的例子中，提交顺序是：sqe5 → sqe2 → sqe3。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br></pre></td><td class="code"><pre><span class="line">SQ array: [5, 2, 3]</span><br><span class="line">SQEs:     [sqe0, sqe1, sqe2, sqe3, sqe4, sqe5]</span><br></pre></td></tr></table></figure><p>在文章中，作者提出一个好处是可以将 request units 放到 internal structure 中，我理解就是后面看到的自定义的 <code>app_sq_ring</code>。另外，也能允许在一个操作中提交多个 sqe。我理解就是如下代码所示，先 fill sqe，再写 array 的操作。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">io_uring_sqe</span> *<span class="title">sqe</span>;</span></span><br><span class="line"><span class="keyword">unsigned</span> tail, index;</span><br><span class="line">tail = sqring-&gt;tail;</span><br><span class="line">index = tail &amp; (*sqring-&gt;ring_mask);</span><br><span class="line">sqe = &amp;sqring-&gt;sqes[index];</span><br><span class="line"><span class="comment">/* this call fills in the sqe entries for this IO */</span></span><br><span class="line">init_io(sqe);</span><br><span class="line"><span class="comment">/* fill the sqe index into the SQ ring array */</span></span><br><span class="line">sqring-&gt;<span class="built_in">array</span>[index] = index;</span><br><span class="line">tail++;</span><br><span class="line">write_barrier();</span><br><span class="line">sqring-&gt;tail = tail;</span><br><span class="line">write_barrier();</span><br></pre></td></tr></table></figure><p>只要 sqe 被内核消费了，application 就可以复用 sqe entry，即使内核还没有完全处理完毕，内核会在需要的时候复制这个结构。</p><p>这样，sqe 的生命周期就比较短，而 application 可能会发送更多的 submission，从而导致 CQ ring 可能溢出。所以默认下的 CQ ring 的大小是 SQ ring 的两倍。</p><p>Completion events 可能以任意顺序到达，它和 submission 的顺序是没有关系的。SQ 和 CQ 两个 ring 是独立运行的。但是每个 submission 事件和每个 completion 事件都能一一对应。</p><h1 id="5-0-io-uring-interface"><a href="#5-0-io-uring-interface" class="headerlink" title="5.0 io_uring interface"></a>5.0 io_uring interface</h1><p>下面介绍的是 io_uring 的“裸”接口，即 system call。</p><h2 id="io-uring-setup"><a href="#io-uring-setup" class="headerlink" title="io_uring_setup"></a>io_uring_setup</h2><p>entries 的取值是 1..=4096，表示 sqe 的数量，必须是 2 的幂。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line"><span class="function"><span class="keyword">int</span> <span class="title">io_uring_setup</span><span class="params">(<span class="keyword">unsigned</span> entries, struct io_uring_params *params)</span></span>;</span><br></pre></td></tr></table></figure><p>params 如下</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br><span class="line">20</span><br><span class="line">21</span><br><span class="line">22</span><br><span class="line">23</span><br><span class="line">24</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">io_uring_params</span> &#123;</span></span><br><span class="line">   <span class="comment">// 由内核填写，表示支持多少个 sqe</span></span><br><span class="line">   __u32 sq_entries;</span><br><span class="line">   <span class="comment">// 由内核填写，表示支持多少个 cqe</span></span><br><span class="line">   __u32 cq_entries;</span><br><span class="line">   __u32 flags;</span><br><span class="line">   __u32 sq_thread_cpu;</span><br><span class="line">   __u32 sq_thread_idle;</span><br><span class="line">   __u32 resv[<span class="number">5</span>];</span><br><span class="line">   <span class="class"><span class="keyword">struct</span> <span class="title">io_sqring_offsets</span> <span class="title">sq_off</span>;</span></span><br><span class="line">   <span class="class"><span class="keyword">struct</span> <span class="title">io_cqring_offsets</span> <span class="title">cq_off</span>;</span></span><br><span class="line">&#125;;</span><br><span class="line"></span><br><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">io_sqring_offsets</span> &#123;</span></span><br><span class="line">   __u32 head; <span class="comment">/* offset of ring head */</span></span><br><span class="line">   __u32 tail; <span class="comment">/* offset of ring tail */</span></span><br><span class="line">   __u32 ring_mask; <span class="comment">/* ring mask value */</span></span><br><span class="line">   __u32 ring_entries; <span class="comment">/* entries in ring */</span></span><br><span class="line">   __u32 flags; <span class="comment">/* ring flags */</span></span><br><span class="line">   __u32 dropped; <span class="comment">/* number of sqes not submitted */</span></span><br><span class="line">   __u32 <span class="built_in">array</span>; <span class="comment">/* sqe index array */</span></span><br><span class="line">   __u32 resv1;</span><br><span class="line">   __u64 resv2;</span><br><span class="line">&#125;;</span><br></pre></td></tr></table></figure><p>io_uring_setup 返回的 int，实际上是一个 fd。如之前所说，这个 fd 是被内核和用户共享的。而 sq_off 和 cq_off 就表示了这共享内存中，SQ 和 CQ 的位置。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br></pre></td><td class="code"><pre><span class="line"><span class="meta">#<span class="meta-keyword">define</span> IORING_OFF_SQ_RING          0ULL</span></span><br><span class="line"><span class="meta">#<span class="meta-keyword">define</span> IORING_OFF_CQ_RING  0x8000000ULL <span class="comment">// 128MB，指的那个 index 数组</span></span></span><br><span class="line"><span class="meta">#<span class="meta-keyword">define</span> IORING_OFF_SQES    0x10000000ULL</span></span><br></pre></td></tr></table></figure><p>用户可以自定义 sq ring 的结构，这个结构中的每个字段都是一个指向到共享内存中位置的指针</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">app_sq_ring</span> &#123;</span></span><br><span class="line">   <span class="keyword">unsigned</span> *head;</span><br><span class="line">   <span class="keyword">unsigned</span> *tail;</span><br><span class="line">   <span class="keyword">unsigned</span> *ring_mask;</span><br><span class="line">   <span class="keyword">unsigned</span> *ring_entries;</span><br><span class="line">   <span class="keyword">unsigned</span> *flags;</span><br><span class="line">   <span class="keyword">unsigned</span> *dropped;</span><br><span class="line">   <span class="keyword">unsigned</span> *<span class="built_in">array</span>;</span><br><span class="line">&#125;;</span><br></pre></td></tr></table></figure><p>如下面的 setup 所示，可以看到自定义的 sring 是如何通过 ptr 和 sq_off 组装起来的。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br></pre></td><td class="code"><pre><span class="line"><span class="function">struct app_sq_ring <span class="title">app_setup_sq_ring</span><span class="params">(<span class="keyword">int</span> ring_fd, struct io_uring_params *p)</span></span></span><br><span class="line"><span class="function"></span>&#123;</span><br><span class="line">   <span class="class"><span class="keyword">struct</span> <span class="title">app_sq_ring</span> <span class="title">sqring</span>;</span></span><br><span class="line">   <span class="keyword">void</span> *ptr;</span><br><span class="line">   ptr = mmap(<span class="literal">NULL</span>, p-&gt;sq_off.<span class="built_in">array</span> + p-&gt;sq_entries * <span class="keyword">sizeof</span>(__u32), PROT_READ | PROT_WRITE, MAP_SHARED | MAP_POPULATE,</span><br><span class="line">   ring_fd, IORING_OFF_SQ_RING);</span><br><span class="line">   sring-&gt;head = ptr + p-&gt;sq_off.head;</span><br><span class="line">   sring-&gt;tail = ptr + p-&gt;sq_off.tail;</span><br><span class="line">   sring-&gt;ring_mask = ptr + p-&gt;sq_off.ring_mask;</span><br><span class="line">   sring-&gt;ring_entries = ptr + p-&gt;sq_off.ring_entries;</span><br><span class="line">   sring-&gt;flags = ptr + p-&gt;sq_off.flags;</span><br><span class="line">   sring-&gt;dropped = ptr + p-&gt;sq_off.dropped;</span><br><span class="line">   sring-&gt;<span class="built_in">array</span> = ptr + p-&gt;sq_off.<span class="built_in">array</span>;</span><br><span class="line">   <span class="keyword">return</span> sring;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><h2 id="io-uring-enter"><a href="#io-uring-enter" class="headerlink" title="io_uring_enter"></a>io_uring_enter</h2><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br></pre></td><td class="code"><pre><span class="line"><span class="function"><span class="keyword">int</span> <span class="title">io_uring_enter</span><span class="params">(</span></span></span><br><span class="line"><span class="function"><span class="params">   <span class="keyword">unsigned</span> <span class="keyword">int</span> fd, <span class="comment">// io_uring_setup 返回的那个 fd</span></span></span></span><br><span class="line"><span class="function"><span class="params">   <span class="keyword">unsigned</span> <span class="keyword">int</span> to_submit, <span class="comment">// tells the kernel that there are up to that amount of sqes ready to be consumed and submitted</span></span></span></span><br><span class="line"><span class="function"><span class="params">   <span class="keyword">unsigned</span> <span class="keyword">int</span> min_complete, <span class="comment">// asks the kernel to wait for completion of that amount of requests.</span></span></span></span><br><span class="line"><span class="function"><span class="params">   <span class="keyword">unsigned</span> <span class="keyword">int</span> flags, </span></span></span><br><span class="line"><span class="function"><span class="params">   <span class="keyword">sigset_t</span> sig)</span></span>;</span><br></pre></td></tr></table></figure><p>可以发现，这个 syscall 可以同时 submit 和 wait for completion，这个也对应了本文作者之前提到的对 aio 的批评之一。</p><p>flags 中有一个参数，设置它，则内核会 actively wait for min_complete events to be available。简单来说，如果希望 wait for completion，则必须设置这个 flag。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line"><span class="meta">#<span class="meta-keyword">define</span> IORING_ENTER_GETEVENTS (1U &lt;&lt; 0)</span></span><br></pre></td></tr></table></figure><h2 id="5-1-SQE-ORDERING"><a href="#5-1-SQE-ORDERING" class="headerlink" title="5.1 SQE ORDERING"></a>5.1 SQE ORDERING</h2><p>这一节主要讲了如何实现 fsync/fdatasync。</p><p>因为之前提到 SQ 和 CQ 是完全独立的，所以这样的机制需要额外的设计。并且因为写入是乱序的，所以我们在乎的是确定所有的写入已经完成。</p><p>io_uring 的机制是，支持 draining the submission side queue，直到之前的 completion 事件都已经结束。在这之前，application 会将后续写入入队。</p><p>通过 IOSQE_IO_DRAIN 这个 flag 来实现这个特性，它会 stall 住整个 SQ。因此，application 可以考虑使用多个 io_uring context，来保证不相关的写是并行的。</p><blockquote><p>io_uring supports draining the submission side queue until all previous completions have finished. This allows the application queue the above mentioned sync operation and know that it will not start before all previous commandshave completed.</p></blockquote><h2 id="5-2-LINKED-SQES"><a href="#5-2-LINKED-SQES" class="headerlink" title="5.2 LINKED SQES"></a>5.2 LINKED SQES</h2><p>所有连续的指定了 IOSQE_IO_LINK 的 io 请求会被串联起来执行，这些请求一定是按照顺序执行的。但是它们和没有指定 IOSQE_IO_LINK 这个 flag 的请求之间的关系是不确定的。</p><h2 id="5-3-TIMEOUT-COMMANDS"><a href="#5-3-TIMEOUT-COMMANDS" class="headerlink" title="5.3 TIMEOUT COMMANDS"></a>5.3 TIMEOUT COMMANDS</h2><h1 id="6-0-Memory-ordering"><a href="#6-0-Memory-ordering" class="headerlink" title="6.0 Memory ordering"></a>6.0 Memory ordering</h1><p>在 <a href="/2017/12/28/Concurrency-Programming-Compare/">并发编程重要概念及比较</a> 中，我们知道 memory order 主要是考虑读-写和写-写问题，如下所示：</p><blockquote><p>read_barrier(): Ensure previous writes are visible before doing subsequent memory reads.<br>write_barrier(): Order this write after previous writes.</p></blockquote><p>我们也知道，不同的 CPU 架构的乱序执行逻辑是不一样的，所以这里只是讨论概念。</p><p>考虑用户侧写入一个 seq，并且通知 kernel 可以去消费了。这就包含了两个过程：</p><ul><li>填写 sqe 中的字段，并且将 sqe index 写入 SQ ring array</li><li>更新 SQ ring 队列的 tail</li></ul><p>这个操作可以简化成下面的伪代码，每一行代表一个内存操作。如果没有合适的 memory order，CPU 是有理由进行乱序执行的。也就是说，无法保证 write 7 是在最后执行的。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br></pre></td><td class="code"><pre><span class="line">1: sqe→opcode = IORING_OP_READV;</span><br><span class="line">2: sqe→fd = fd;</span><br><span class="line">3: sqe→off = 0;</span><br><span class="line">4: sqe→addr = &amp;iovec;</span><br><span class="line">5: sqe→len = 1;</span><br><span class="line">6: sqe→user_data = some_value;</span><br><span class="line">7: sqring→tail = sqring→tail + 1;</span><br></pre></td></tr></table></figure><p>所以，需要添加如下的 write barrier</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br></pre></td><td class="code"><pre><span class="line">1: sqe→opcode = IORING_OP_READV;</span><br><span class="line">2: sqe→fd = fd;</span><br><span class="line">3: sqe→off = 0;</span><br><span class="line">4: sqe→addr = &amp;iovec;</span><br><span class="line">5: sqe→len = 1;</span><br><span class="line">6: sqe→user_data = some_value;</span><br><span class="line"> write_barrier(); /* ensure previous writes are seen before tail write */</span><br><span class="line">7: sqring→tail = sqring→tail + 1;</span><br><span class="line"> write_barrier(); /* ensure tail write is seen */</span><br></pre></td></tr></table></figure><h1 id="7-0-liburing-library"><a href="#7-0-liburing-library" class="headerlink" title="7.0 liburing library"></a>7.0 liburing library</h1><p>通过这个库，可以：</p><ul><li>不需要写一堆 boiler plate code</li><li>不需要考虑 memory ordering 的问题</li><li>不需要考虑自己维护 ring buffer 的问题</li></ul><h2 id="7-1-LIBURING-IO-URING-SETUP"><a href="#7-1-LIBURING-IO-URING-SETUP" class="headerlink" title="7.1 LIBURING IO_URING SETUP"></a>7.1 LIBURING IO_URING SETUP</h2><h1 id="8-0-Advanced-use-cases-and-features"><a href="#8-0-Advanced-use-cases-and-features" class="headerlink" title="8.0 Advanced use cases and features"></a>8.0 Advanced use cases and features</h1><h2 id="8-1-FIXED-FILES-AND-BUFFERS"><a href="#8-1-FIXED-FILES-AND-BUFFERS" class="headerlink" title="8.1 FIXED FILES AND BUFFERS"></a>8.1 FIXED FILES AND BUFFERS</h2><h2 id="8-2-POLLED-IO"><a href="#8-2-POLLED-IO" class="headerlink" title="8.2 POLLED IO"></a>8.2 POLLED IO</h2><h2 id="8-3-KERNEL-SIDE-POLLING"><a href="#8-3-KERNEL-SIDE-POLLING" class="headerlink" title="8.3 KERNEL SIDE POLLING"></a>8.3 KERNEL SIDE POLLING</h2><h1 id="9-0-Performance"><a href="#9-0-Performance" class="headerlink" title="9.0 Performance"></a>9.0 Performance</h1><h1 id="Reference"><a href="#Reference" class="headerlink" title="Reference"></a>Reference</h1><ul><li><a href="https://man7.org/linux/man-pages/man2/io_submit.2.html" target="_blank" rel="noopener">https://man7.org/linux/man-pages/man2/io_submit.2.html</a></li></ul>]]></content>
    
    
    <summary type="html">&lt;p&gt;通过 &lt;a href=&quot;https://kernel.dk/io_uring.pdf&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://kernel.dk/io_uring.pdf&lt;/a&gt; 简单学习下 io_uring。&lt;/p&gt;</summary>
    
    
    
    
    <category term="Linux" scheme="http://www.calvinneo.com/tags/Linux/"/>
    
    <category term="FileSystem" scheme="http://www.calvinneo.com/tags/FileSystem/"/>
    
  </entry>
  
  <entry>
    <title>LLM 基础概念和核心问题整理</title>
    <link href="http://www.calvinneo.com/2025/10/25/on-llm/"/>
    <id>http://www.calvinneo.com/2025/10/25/on-llm/</id>
    <published>2025-10-24T18:20:13.000Z</published>
    <updated>2026-05-26T10:55:18.047Z</updated>
    
    <content type="html"><![CDATA[<p>主要关注：</p><ul><li>LLM 的基础原理</li><li>Transformer 的架构</li><li>KVCache</li></ul><p>基本来自于我的问题，以及我对通过 chat 学习到的东西的总结，由 AI 整理。完整版我放在腾讯文档中，这里选取一些方便摘录的写成博客。</p><a id="more"></a><h1 id="LLM-的基础理论-AI-Infra"><a href="#LLM-的基础理论-AI-Infra" class="headerlink" title="LLM 的基础理论(AI Infra)"></a>LLM 的基础理论(AI Infra)</h1><h2 id="大模型的计算复杂度"><a href="#大模型的计算复杂度" class="headerlink" title="大模型的计算复杂度"></a>大模型的计算复杂度</h2><h3 id="Prefill-和-Decode"><a href="#Prefill-和-Decode" class="headerlink" title="Prefill 和 Decode"></a>Prefill 和 Decode</h3><p>N 为上下文长度：</p><ul><li>如果没有 KVCache，O(N^2)</li><li>如果有 KVCache，O(N)</li></ul><p>O(N) 的原因：</p><ul><li>Prefill 阶段，需要拿一个 $N \times d$ 的 Query 矩阵，去乘以一个 $d \times N$ 的 Key 矩阵（$d$ 是维度）。这会生成一个 $N \times N$ 的注意力分数矩阵（Attention Map）。这个对应了“为什么当你扔进去一本 10 万字的小说，模型在吐出第一个字之前，会卡顿很久（首字延迟 TTFT 很高）。”<ul><li>当然，这里和 N_head 也有关系。每个 head 是单独算的。</li></ul></li><li>Decode 阶段，这里是 KVCache 的优化了。<ul><li>没有 KV Cache 的傻瓜做法： 假设模型要吐出第 $N+1$ 个词。如果没缓存，它必须把前面的 $N$ 个词加上新词，重新过一遍 $O((N+1)^2)$ 的计算。这显然极其愚蠢且缓慢。</li><li>有了 KV Cache 的聪明做法： 因为前 $N$ 个词的 Key 和 Value 已经存在显存里了。对于第 $N+1$ 个词，模型只需要单独为这 1 个新词计算它的 Query（大小为 $1 \times d$）。</li><li>数学本质： 此时的注意力计算，变成了拿这个 $1 \times d$ 的 Query 向量，去乘以缓存中 $d \times N$ 的 Key 矩阵。</li><li>复杂度： 向量乘矩阵，计算复杂度瞬间从 $O(N^2)$ 暴降到了 $O(N)$！</li></ul></li></ul><p>两个耗时：</p><ul><li><p>TTFT (Time To First Token - 首字时间)： 衡量 Prefill（预填充） 阶段的耗时。即你按下回车，到屏幕上蹦出第一个字等了多久。</p><ul><li>优势： 它是高度并行的。现代 GPU 非常擅长干这种巨大的矩阵乘法，虽然计算量大，但 GPU 的几万个核心都在同时工作，效率极高。这被称为 Compute-bound（算力瓶颈）。</li></ul></li><li><p>TPOT (Time Per Output Token - 每输出词元时间)： 衡量 Decode（解码） 阶段的速度。即后续的字像吐牙膏一样，每个字蹦出来平均要多久。</p><ul><li>为什么耗时： 它必须是串行（Sequential）的！因为模型必须等第 1 个字出来，才能去算第 2 个字。在这个阶段，计算量虽然降到了 $O(N)$，但就像我们之前说的，为了算这 1 个新词，GPU 必须把显存里庞大的 KV Cache 重新搬运一遍到计算单元。</li><li>劣势： 此时 GPU 的算力核心其实大部分时间在“闲置摸鱼”，它在等显存（HBM）把数据一点点搬过来。这被称为 Memory-bound（内存带宽瓶颈）。</li></ul></li><li><p>场景 A：长输入，短输出（Prefill 耗时长）</p></li><li><p>场景 B：短输入，长输出（Decode 耗时长）</p></li></ul><p>尽管如此，O(N) 依然是恐怖的：<br>当 $N$ 达到 1,000,000 时，即便只是单步 $O(N)$ 的计算，对于内存带宽的压力也是极其巨大的。因为你要为了这一个新词，把显存里 100 万个历史 Key 和 Value 全都搬运到计算单元里比对一次。这种现象叫做 Memory Bound（内存带宽瓶颈），此时 GPU 算力可能只发挥了 5%，剩下的时间都在等显存传输数据。</p><h3 id="关于注意力头"><a href="#关于注意力头" class="headerlink" title="关于注意力头"></a>关于注意力头</h3><p>在阅读大模型论文或源码时，经常会看到两个容易混用的“维度”概念：</p><ul><li>模型的总维度 ($d_{\text{model}}$)： 这是模型在每一层之间传递数据的整体“宽度”。比如一个典型的 13B 模型，它的 $d_{\text{model}}$ 可能是 5120。</li><li>注意力头的维度 ($d_{\text{head}}$)： 因为大模型采用的是“多头机制（Multi-Head）”，它不会拿 5120 这么庞大的向量直接去算，而是把它“切分”给多个不同的“头”去并行计算，让每个头关注不同的特征。</li></ul><p>它们的严格数学关系是：<br>$$<br>d_{\text{model}} = N_{\text{heads}} \times d_{\text{head}}<br>$$</p><p>例如，把 5120 的总维度分给 40 个头（$N_{\text{heads}} = 40$），那么每个头的维度 $d_{\text{head}}$ 就是 128。</p><h2 id="PagedAttention"><a href="#PagedAttention" class="headerlink" title="PagedAttention"></a>PagedAttention</h2><p>vLLM 是一个用于大语言模型（LLM）的开源高性能推理和服务引擎。vLLM 最具革命性的特性是 PagedAttention 技术：传统方法分配 KV Cache 像住宾馆，必须预留连续的、足够大的房间（连续显存），导致大量显存碎片被浪费。PagedAttention 借鉴了操作系统的虚拟内存机制，将 KV Cache 拆分成一个个极小的“内存页”（Block），按需分配，非连续存放。这几乎消灭了显存碎片，让系统能同时服务的用户数翻了好几倍。</p><p>架构层面：</p><ul><li>vLLM 就像是生产线上的“超级组装机”，负责飞速地计算和吐出 Token</li><li>3FS 则是紧挨着生产线的“超高速立体仓库”，专门负责存放那些流水线上放不下的半成品（KV Cache）。<br>  vLLM 本身自带了前缀缓存（Prefix Caching）功能，但这通常局限在单台服务器内部。3FS 的目的是全局视角的缓存共享。</li></ul><p>PagedAttention 是如何实现的，如下其实还是比较传统的：</p><ul><li>它将每个用户请求的 KV Cache 划分成固定大小的“块”（Block）。每个块不再是以字节（Byte）为单位，而是以 Token 数量为单位（默认通常是 16 个或 32 个 Token）。</li><li>逻辑与物理分离 + 维护块表（Block Table）</li></ul><p>不同点：</p><ul><li>OS 是以 Bytes 为基础。PA 以 Token 为基础。一个 Block 占用多少显存，取决于模型的层数、注意力头数和数据精度。例如，在 Llama-2 13B 模型中，一个 16-Token 的 Block 可能占用几兆字节（MB）的显存。</li><li>访问模式不同：OS 是随机访问多。PA 是可预测的，大模型的文本生成（自回归）是严格按顺序追加（Append-only）的。模型永远是一个词一个词往后吐，KV Cache 只会顺着时间轴往后写。</li><li>触发 offloading 的时机不同：OS 就是缺页中断。vLLM 是在<strong>应用层（框架层）</strong>预测显存够不够。如果不够了，vLLM 会主动把一些 Block 踢到系统内存（Offload），这被称为抢占（Preemption）或重计算（Recomputation），完全由软件逻辑控制。</li></ul><p>那么如何利用这个特性？</p><ul><li>减少碎片<ul><li>消灭内部碎片： 一个物理块一旦分配，模型就会按顺序把它填满。除了当前正在生成的最后那个未填满的 Block，所有已分配的 Block 的空间利用率都是 100%。</li><li>消灭外部碎片： 因为物理块在显存中是不连续的，只要显存里还有任何一个空闲的 Block，就可以立刻分配给任何请求。这让显存的整体利用率从传统方案的 20%-40% 瞬间飙升到了 90% 以上。</li></ul></li><li>由于 PagedAttention 是纯软件层面的控制，它拥有“全局视野”。所以它可以选择重新计算，还是 offload 到 3FS 中。</li></ul><h3 id="Hybrid-Sparse-Attention"><a href="#Hybrid-Sparse-Attention" class="headerlink" title="Hybrid Sparse Attention"></a>Hybrid Sparse Attention</h3><p>DeepSeek（特别是在其最新的 V4 架构中）采用的 Hybrid Sparse Attention，是从算法数学层面下刀，它混合了多种注意力策略来“偷懒”却不掉精度：</p><ul><li>滑动窗口注意力 (Sliding Window Attention, SWA)： 对最近的一小段上下文（比如最近 128 个 Token），采用全量密集注意力，保证局部的绝对精确。</li><li>压缩稀疏注意力 (Compressed Sparse Attention, CSA)： 对远古的历史上下文，它不再让当前 Token 和所有历史 Token 一一比对，而是利用轻量级的“索引器（Indexer）”挑选出最相关的 Top-K 个 Token，只计算这部分极其重要的记忆，其余的直接忽略（稀疏化）。</li><li>重度压缩注意力 (Heavily Compressed Attention, HCA)： 把极其冗长的历史记录，像打“压缩包”一样，将多个 Token 融合压缩成少量的特征向量，提供一个全局视角的模糊记忆。</li></ul><h2 id="vLLM-和训推存储底座"><a href="#vLLM-和训推存储底座" class="headerlink" title="vLLM 和训推存储底座"></a>vLLM 和训推存储底座</h2><p>vLLM 并没有完全撒手不管“存”。它管理最宝贵、最快的 GPU 显存（HBM）。一旦发现显存快满了，就立刻把暂时不用的 KV Cache 通过超高速网络（如 RDMA）卸载给 3FS。<br>RDMA（Remote Direct Memory Access，远程直接内存访问）是一种无需计算机操作系统内核接入、无需CPU参与，即可直接在两台机器内存之间读写数据的网络技术。它通过内核旁路（Kernel Bypass）和零拷贝技术，实现了高吞吐量、低延迟和低CPU占用的网络传输，适用于高性能计算和大数据中心。<br>RDMA 编程通常基于 Verbs API (如 libibverbs) 进行，核心流程是建立连接并交换内存地址信息，随后进行直接读写。 </p><ul><li>内存注册 (Memory Registration, MR)： RDMA 必须先将应用内存区域（Buffer）注册到网卡，锁定物理地址，并获取远程访问的键（R_Key）。</li><li>队列对 (Queue Pair, QP)： 每个RDMA应用通过QP通信，包括发送队列（SQ）和接收队列（RQ）。</li><li>工作请求 (Work Request, WR)： 应用通过WR提交发送/接收指令给网卡。</li><li>完成队列 (Completion Queue, CQ)： 网卡完成传输后，将完成结果放入 CQ，应用通过轮询 CQ 获取结果。</li></ul><p>这套机制的作用：</p><ul><li>独立扩容：vLLM 加显卡，3FS 加 SSD</li><li>无状态的 vLLM</li><li>多级缓存：VRAM（显存） -&gt; DRAM（内存） -&gt; 3FS（高速网络 SSD）</li></ul><h2 id="前缀"><a href="#前缀" class="headerlink" title="前缀"></a>前缀</h2><h3 id="关于前缀"><a href="#关于前缀" class="headerlink" title="关于前缀"></a>关于前缀</h3><ol><li>最佳实践：调整 Prompt 结构（把“死”的放前面）<ol><li>失效做法： [用户闲聊] + [系统指令] + [长财报]（用户一换，缓存全崩）。</li><li>优化做法： [长财报] + [系统指令] + [用户闲聊]。</li></ol></li><li>基数树（Radix Tree）管理机制<ol><li>如果你的输入是 A + B（A 是闲聊，B 是财报），它会先找 A 的缓存。</li><li>虽然它发现 A + B 不能直接复用 C + B 的缓存，但如果系统发现大量请求都有共同的 B，它会尝试在内部进行“多级缓存”匹配。</li></ol></li><li>针对位置编码的魔改：相对位置编码（RoPE）与偏移</li><li>DeepSeek 的特殊方案：跨序列的注意力（Multi-Head Latent Attention）</li></ol><h2 id="距离和位置编码"><a href="#距离和位置编码" class="headerlink" title="距离和位置编码"></a>距离和位置编码</h2><p>LLM 对“距离”的感知是弱结构化、统计性的，而不是显式离散结构化的：</p><ol><li><p>Transformer 本质上是“全连接注意力”</p><ol><li>任意 token 都可以直接看到任意 token</li><li>不像 RNN 那样必须一步一步传递</li><li>所以模型天然没有“近的更近、远的更远”这种 inductive bias</li><li>如下所示，A attention 到 B 和 E 的距离是一样的 <figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">A B C D E</span><br></pre></td></tr></table></figure></li></ol></li><li><p>Position Embedding 只是“补充位置信息”<br> 因为 attention 本身没有顺序概念，所以必须人为加 position encoding。诸如 <code>token_embedding + position_embedding</code> 或者 RoPE 等都是一种形式。</p></li><li><p>它“知道顺序”，但不是算法式知道</p> <figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">Tom gave Jerry a book because he trusted him.</span><br></pre></td></tr></table></figure><p> 大模型能够语义地知道 Tom 这里是 he。但是它不是通过 Tom – he，Jerry – him 这样的位置对应关系知道的。</p></li></ol><p>随着上下文变长：</p><ul><li>nesting depth 增大</li><li>distribution shift 出现</li><li>模型就容易崩</li></ul><p>这也是下面这些比较难做的原因：</p><ul><li>long-context reasoning</li><li>counting</li><li>exact matching</li></ul><h3 id="Lost-in-the-Middle-问题"><a href="#Lost-in-the-Middle-问题" class="headerlink" title="Lost in the Middle 问题"></a>Lost in the Middle 问题</h3><p>指的是 Transformer 对“长距离结构”和“中间位置”缺乏稳定位置感的一种表现。</p><h1 id="Transformer"><a href="#Transformer" class="headerlink" title="Transformer"></a>Transformer</h1><h2 id="Attention-机制"><a href="#Attention-机制" class="headerlink" title="Attention 机制"></a>Attention 机制</h2><p>NLP 对 token 序列 X 的三种编码方式：</p><ul><li><p>RNN<br>  RNN 是递归的结构，所以只能串行计算。</p>  <figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">y_t = f(y_&#123;t-1&#125;, x_t)</span><br></pre></td></tr></table></figure></li><li><p>CNN<br>  CNN 能够并行计算，但是因为引入了窗口，所以只能看到局部信息。</p>  <figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">y_t = f(x_&#123;t-1&#125;, x_t, x_&#123;t+1&#125;)</span><br></pre></td></tr></table></figure></li><li><p>Attention</p></li></ul><h3 id="Q、K、V"><a href="#Q、K、V" class="headerlink" title="Q、K、V"></a>Q、K、V</h3><p>从定义上看，对于 token 流</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">x1, x2, x3, ... xn</span><br></pre></td></tr></table></figure><p>每个 xi 通过三组线性变换生成：</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br></pre></td><td class="code"><pre><span class="line">Qi = xi * Wq</span><br><span class="line">Ki = xi * Wk</span><br><span class="line">Vi = xi * Wv</span><br></pre></td></tr></table></figure><p>考虑下面的句子</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">The animal didn’t cross the street because it was too tired.</span><br></pre></td></tr></table></figure><p>句子中的每一个 token，都有一个自己的 Q。用 <code>Wq</code> 可以提取出这个 Q，如下：</p><ul><li>animal → 它在问：后面有没有补充信息？</li><li>because → 它在问：因果关系是什么？</li><li>it → 它在问：我指的是谁？</li><li>tired → 它在问：谁在 tired？</li></ul><p>对于 K，则告诉了这个 token 可以回答什么样的 Q：</p><ul><li>animal → 我是一个名词、可能是指代目标</li><li>street → 我是一个地点名词</li><li>because → 我是因果连接词</li><li>tired → 我是状态形容词</li><li>cross → 我是动作</li></ul><p>对于 V，承载了语义信息的本体：</p><ul><li>animal → 动物这个实体的语义</li><li>street → 街道的概念</li><li>tired → 疲劳的状态语义</li><li>cross → 穿越动作</li></ul><h3 id="Self-Attention"><a href="#Self-Attention" class="headerlink" title="Self-Attention"></a>Self-Attention</h3><p>Self-Attention 指的是 Q、K、V 都来自同一组 token 的 attention。即</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br></pre></td><td class="code"><pre><span class="line">X → Q = XWq</span><br><span class="line">X → K = XWk</span><br><span class="line">X → V = XWv</span><br></pre></td></tr></table></figure><p>然后</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">Attention(Q, K, V)</span><br></pre></td></tr></table></figure><p>为什么叫 Self？</p><ul><li>不是去外部文档查</li><li>不是去另一段文本查</li><li>而是在自己这段序列内部相互对齐</li></ul><p>对应的是 Cross-Attention：</p><ul><li>Q 来自 Decoder</li><li>K/V 来自 Encoder</li></ul><p>容易发现，Cross-Attention 更适合翻译或者对齐另一段文本。而 Self-Attention 更适合理解一句话内部关系。</p><h3 id="Multi-Head-Attention"><a href="#Multi-Head-Attention" class="headerlink" title="Multi-Head Attention"></a>Multi-Head Attention</h3><p>如果一个 token 同时想问多种不同类型的问题怎么办？引入多组问题呗。所以上面的三个矩阵会变成三组矩阵。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br></pre></td><td class="code"><pre><span class="line">Wq¹ Wk¹ Wv¹</span><br><span class="line">Wq² Wk² Wv²</span><br><span class="line">Wq³ Wk³ Wv³</span><br><span class="line">...</span><br></pre></td></tr></table></figure><p>还是对上面的例子而言</p><p>对 it 这个 token：</p><ul><li>Head1 问“我指代谁？”，会关注 animal（指代）</li><li>Head2 问“是否有因果关系？”，会关注 because（因果）</li><li>Head3 问“我与哪个动词相关？”，会关注 cross（动作）</li></ul><h2 id="KVCache"><a href="#KVCache" class="headerlink" title="KVCache"></a>KVCache</h2><h3 id="Why"><a href="#Why" class="headerlink" title="Why"></a>Why</h3><p>Token 是模型处理文本的最小离散单位。所以 LLM 并不是直接处理文字，而是直接处理 token。Token 是通过分词器从文本切出来的子串单位。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">&quot;Hello world&quot; → [15496, 2159]</span><br></pre></td></tr></table></figure><p>但是</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br></pre></td><td class="code"><pre><span class="line">&quot;refund&quot; = 1 token</span><br><span class="line">&quot;refunding&quot; = [&quot;refund&quot;, &quot;ing&quot;]</span><br></pre></td></tr></table></figure><p>不同的模型能够接受不同的上下文长度，因此，它们的 KVCache 也要更大：</p><ul><li>GPT-3.5 4k tokens</li><li>GPT-4   8k / 32k</li><li>Claude  100k</li></ul><p>那么 token 是不是特指用户 prompt输入的 token 或者模型输出的 token 呢？其实根据下面的自回归生成，这两个是一个东西。</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">prompt1 → prompt2 → prompt3 → output1 → output2 ...</span><br></pre></td></tr></table></figure><p>自回归生成:模型按顺序生成 token，每个 token 都只依赖之前已经生成的 token。<br>例如，下面的 token 序列中，x1 到 x3 是用户的 prompt 输入</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">x1 → x2 → x3 → x4 → ... → xT</span><br></pre></td></tr></table></figure><p>则 xT 生成的方式是</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">P(x_t | x_1, x_2, ..., x_&#123;T-1&#125;)</span><br></pre></td></tr></table></figure><p>由此可见，自回归生成导致推理是串行的。为什么要自回归生成呢？原因比较深入，可以理解为：</p><ul><li>语言本质是序列，唯一通用可行的分解方式就是 chain rule</li><li>自回归训练极其稳定<br>  输入是前缀，目标是预测下一个 token，loss 是交叉熵。不需要强化学习。</li><li>从历史演化角度来看，n-gram、RNN、LSTM 等都是自回归的</li></ul><p>因为自回归生成，所以导致了 KVCache 的出现。KVCache 把 O(n²) 变成 O(n)。</p><p>KVCache 具有如下的形式：</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">KVCache[layer][token_index] = (K, V)</span><br></pre></td></tr></table></figure><ul><li><p>layer 是 Transformer 架构的层<br>  现在的 LLM 大都是 Decoder only 的架构，所以这里的层如下所示。</p>  <figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br></pre></td><td class="code"><pre><span class="line">Input tokens</span><br><span class="line">  ↓</span><br><span class="line">Embedding</span><br><span class="line">  ↓</span><br><span class="line">Decoder Block 1</span><br><span class="line">  ↓</span><br><span class="line">Decoder Block 2</span><br><span class="line">  ↓</span><br><span class="line">...</span><br><span class="line">  ↓</span><br><span class="line">Decoder Block N</span><br><span class="line">  ↓</span><br><span class="line">LM Head</span><br></pre></td></tr></table></figure><p>  不同层的 schema 都不同：</p><ul><li>第一层：按字形索引</li><li>中间层：按语法索引</li><li>高层：按语义索引</li></ul><p>  每一层内部包含：</p><ul><li>Masked Self-Attention<br>  让每个 token 从历史 token 中选择性地读取信息。<br>  当前 token 只能看到自己和之前的 token，不能看到未来的 token。</li><li>FFN<br>  这里就是前馈神经网络，目的是在单个 token 维度上做语义升维和重映射。可以看成是在理解输入的 token。</li><li>Residual / Norm</li></ul></li><li><p>token_index 表示这是第几个 token<br>  因为 attention 在第 t 步需要：当前 token 的 Q、对比所有历史 token 的 K、加权读取所有历史 token 的 V。所以这些 K 和 V 需要按照 token_index 来存储。</p></li><li><p>K 和 V<br>  K 是我提供什么信息给别人关注。<br>  V 是别人关注我时能读到什么内容。</p></li><li><p>为什么不需要缓存 Q？<br>  因为 Q 只在当前步骤被使用。<br>  在第 t 步，会用 <code>Q_t</code> 去访问 <code>K_{0..t-1}</code>, <code>V_{0..t-1}</code>。但是在未来，不会再去访问 <code>Q_t</code> 了。</p></li></ul><p>如果没有 KVCache，每生成一个新 token 需要重新计算所有历史 token 的 K/V 复杂度是 O(n²)</p><h1 id="Reference"><a href="#Reference" class="headerlink" title="Reference"></a>Reference</h1><ul><li><a href="https://zh.d2l.ai/chapter_attention-mechanisms/attention-cues.html" target="_blank" rel="noopener">https://zh.d2l.ai/chapter_attention-mechanisms/attention-cues.html</a></li><li><a href="https://transformers.run/c1/attention/" target="_blank" rel="noopener">https://transformers.run/c1/attention/</a></li><li><a href="https://zhuanlan.zhihu.com/p/338817680" target="_blank" rel="noopener">https://zhuanlan.zhihu.com/p/338817680</a></li></ul>]]></content>
    
    
    <summary type="html">&lt;p&gt;主要关注：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LLM 的基础原理&lt;/li&gt;
&lt;li&gt;Transformer 的架构&lt;/li&gt;
&lt;li&gt;KVCache&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;基本来自于我的问题，以及我对通过 chat 学习到的东西的总结，由 AI 整理。完整版我放在腾讯文档中，这里选取一些方便摘录的写成博客。&lt;/p&gt;</summary>
    
    
    
    
    <category term="机器学习" scheme="http://www.calvinneo.com/tags/机器学习/"/>
    
    <category term="LLM" scheme="http://www.calvinneo.com/tags/LLM/"/>
    
  </entry>
  
  <entry>
    <title>Persistent data structures</title>
    <link href="http://www.calvinneo.com/2025/10/06/persistent-data-structures/"/>
    <id>http://www.calvinneo.com/2025/10/06/persistent-data-structures/</id>
    <published>2025-10-05T18:20:13.000Z</published>
    <updated>2025-11-21T17:13:07.931Z</updated>
    
    <content type="html"><![CDATA[<p>在 rust 中，immutable 的数据结构的性质是非常好的。在大部分函数式语言中，都不允许存在 mutable 的数据。</p><p>如果要在不可变数据结构上进行修改，就需要 clone 一份出来。因此：</p><ul><li>对于一些较大的结构，希望能够尽量复用</li><li>如果此时只有一份引用，则可以直接获取 mut 引用就地修改</li></ul><p>所以有了 Persistent data structures 的概念：</p><ul><li>每一次修改该结构，都会保留之前的版本</li><li>历史的版本可以被查询</li><li>如果历史版本的数据也支持修改，则称为 Full persistence，否则称为 Partial persistence</li></ul><a id="more"></a><h1 id="实现方案"><a href="#实现方案" class="headerlink" title="实现方案"></a>实现方案</h1><h2 id="Copy-on-Write"><a href="#Copy-on-Write" class="headerlink" title="Copy on Write"></a>Copy on Write</h2><p>用一个数组存放所有的历史版本，非常 bruteforce。</p><h2 id="Fat-node"><a href="#Fat-node" class="headerlink" title="Fat node"></a>Fat node</h2><p>为每一个 field 维护历史记录。例如</p><figure class="highlight c++"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">Node</span> &#123;</span></span><br><span class="line">    <span class="keyword">int</span> value;</span><br><span class="line">    Node* left;</span><br><span class="line">    Node* right;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>会变成</p><figure class="highlight c++"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">FatNode</span> &#123;</span></span><br><span class="line">    <span class="built_in">vector</span>&lt;Pair&lt;version, value&gt;&gt; value_history;</span><br><span class="line">    <span class="built_in">vector</span>&lt;Pair&lt;version, Node*&gt;&gt; left_history;</span><br><span class="line">    <span class="built_in">vector</span>&lt;Pair&lt;version, Node*&gt;&gt; right_history;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>查询的过程就有点像是 MVCC 了。给定一个 version，去 lower_bound 找到最大的小于 version 的修改。</p><p>容易看到，Fat node 不支持 Full persistence。</p><h2 id="Split"><a href="#Split" class="headerlink" title="Split"></a>Split</h2><p>Fat Node 不能无限制增长，否则：</p><ul><li>历史太长</li><li>查询复杂度上升</li><li>节点缓存局部性变差</li></ul><p>因此，需要将节点 split 为两个节点：</p><ul><li>新节点记录最新版本的值</li><li>老节点保留早期历史</li><li>结构中指向该节点的指针，也在相应版本中被更新为指向新的节点</li></ul><h2 id="Path-copying"><a href="#Path-copying" class="headerlink" title="Path copying"></a>Path copying</h2><h1 id="常见的结构实现"><a href="#常见的结构实现" class="headerlink" title="常见的结构实现"></a>常见的结构实现</h1><h2 id="List"><a href="#List" class="headerlink" title="List"></a>List</h2><h2 id="Vec"><a href="#Vec" class="headerlink" title="Vec"></a>Vec</h2><h1 id="Reference"><a href="#Reference" class="headerlink" title="Reference"></a>Reference</h1><ul><li><a href="https://github.com/orium/rpds" target="_blank" rel="noopener">https://github.com/orium/rpds</a></li></ul>]]></content>
    
    
    <summary type="html">&lt;p&gt;在 rust 中，immutable 的数据结构的性质是非常好的。在大部分函数式语言中，都不允许存在 mutable 的数据。&lt;/p&gt;
&lt;p&gt;如果要在不可变数据结构上进行修改，就需要 clone 一份出来。因此：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;对于一些较大的结构，希望能够尽量复用&lt;/li&gt;
&lt;li&gt;如果此时只有一份引用，则可以直接获取 mut 引用就地修改&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;所以有了 Persistent data structures 的概念：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;每一次修改该结构，都会保留之前的版本&lt;/li&gt;
&lt;li&gt;历史的版本可以被查询&lt;/li&gt;
&lt;li&gt;如果历史版本的数据也支持修改，则称为 Full persistence，否则称为 Partial persistence&lt;/li&gt;
&lt;/ul&gt;</summary>
    
    
    
    
    <category term="Rust" scheme="http://www.calvinneo.com/tags/Rust/"/>
    
    <category term="数据结构" scheme="http://www.calvinneo.com/tags/数据结构/"/>
    
  </entry>
  
  <entry>
    <title>Zero-Copy 技术</title>
    <link href="http://www.calvinneo.com/2025/09/30/zero-copy/"/>
    <id>http://www.calvinneo.com/2025/09/30/zero-copy/</id>
    <published>2025-09-30T15:07:22.000Z</published>
    <updated>2026-01-14T18:35:35.431Z</updated>
    
    <content type="html"><![CDATA[<p>介绍 Linux 中的零拷贝技术。从 <a href="/2025/03/09/learn-fuse/">Fuse 学习</a> 中独立出来。</p><a id="more"></a><h1 id="read、write-接口"><a href="#read、write-接口" class="headerlink" title="read、write 接口"></a>read、write 接口</h1><p>从普通文件 read，涉及两次复制：</p><ul><li>从磁盘通过 DMA 读到内核的 page cache<br>  这里的 page cache 机制也是一种 kernel buffer，但专门提供给磁盘文件的。</li><li>从内核的 page cache 复制到 user buffer</li></ul><p>从套接口读数据：</p><ul><li>从网卡通过 DMA 直接写入 kernel buffer</li><li>从 kernel buffer 复制到 user buffer</li></ul><p>注意，在使用 DMA 之前，磁盘读出来的数据会放到一个寄存器里面，然后通过中断通知 CPU 把数据写到临时的内存中攒批，最后写到 page cache 中。但是该方式性能太差，早已经淘汰了。</p><p>读数据过程：</p><ul><li>调用 read() 函数陷入内核，第一次 context switch</li><li>DMA 控制器将数据从磁盘拷贝到 kernel buffer，这是第一次 DMA 拷贝</li><li>CPU 将数据从 kernel buffer 复制到 user buffer，这是第一次 CPU 拷贝</li><li>CPU 完成拷贝之后，read() 函数返回到用户态，第二次 context switch</li></ul><p>写过程类似。</p><h1 id="mmap"><a href="#mmap" class="headerlink" title="mmap"></a>mmap</h1><p>把 kernel space 的页映射到 user space，所以可以避免从 kernel space 到 user space 的一次复制。<br>关于 mmap 可以见 <a href="/2025/01/03/memory-context-knowledge/">内存领域知识</a>。</p><h1 id="sendfile"><a href="#sendfile" class="headerlink" title="sendfile"></a>sendfile</h1><h2 id="原始-sendfile"><a href="#原始-sendfile" class="headerlink" title="原始 sendfile"></a>原始 sendfile</h2><p>sendfile 将数据从磁盘读到内核的 page cache，然后将 page cache 复制到 socket 的 buffer 中。</p><p>它的好处是减少了 syscall 的次数。将 read + write 或者 mmap + write 打包了。<br>但是，仍然需要 2 次 DMA 拷贝和 1 次 CPU 拷贝。</p><h2 id="sendfile-DMA-优化"><a href="#sendfile-DMA-优化" class="headerlink" title="sendfile + DMA 优化"></a>sendfile + DMA 优化</h2><p>将从 page cache 到 socket buffer 的那一次 CPU 拷贝去掉了。DMA 可以直接从 page cache 拷贝数据到网卡里面。</p><h1 id="splice"><a href="#splice" class="headerlink" title="splice"></a>splice</h1><p>限制是 fd_in 和 fd_out 中，至少有一个是 pipe：</p><ul><li>如果 fd_in 是 pipe，那么 off_in 必须是 NULL</li><li>如果 fd_in 不是 pipe，且 off_in 是 NULL，那么 bytes are read from fd_in starting from the file offset, and the file offset is adjusted appropriately.</li><li>如果 fd_in 不是 pipe，且 off_in 不是 NULL，off_in must point to a buffer which specifies the starting offset from which bytes will be read from fd_in; in this case, the file offset of fd_in is not changed, and the offset pointed to by off_in is adjusted appropriately instead.</li></ul><p>这里解释一下什么是 linux 中的管道：</p><ul><li>匿名管道（anonymous pipe）<br>  由父进程创建，用在具有亲缘关系的进程之间通信。<br>  通过 pipe() 系统调用创建，返回一对文件描述符：一个用于写，一个用于读。<br>  只存在于内存中，它不是一个磁盘上的文件，不能用 ls 查看，也没有 inode 号。</li><li>命名管道（named pipe，也叫 FIFO）<br>  具有名字的管道，可以存在于文件系统中，有路径。文件类型是 p，代表 pipe。<br>  通过 mkfifo 命令或者 mkfifo() 系统调用创建。<br>  可以实现非亲缘进程之间的通信。</li></ul><p>所有的匿名管道都支持 splice，通常借助匿名管道来实现 zero copy。此时，pipefd 就起到了中转管道的作用，它连接了两个彼此之间不支持零拷贝的 fd。我觉得是一个比较有意思的设计，通过匿名管道的中介，减少了不同 fd 之间实现相互 zero copy 的复杂度。</p><figure class="highlight c++"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br></pre></td><td class="code"><pre><span class="line">splice(file_fd, <span class="literal">NULL</span>, pipefd[<span class="number">1</span>], <span class="literal">NULL</span>, len, <span class="number">0</span>);</span><br><span class="line">splice(pipefd[<span class="number">0</span>], <span class="literal">NULL</span>, socket_fd, <span class="literal">NULL</span>, len, <span class="number">0</span>);</span><br></pre></td></tr></table></figure><p>一些命名管道也支持 splice，但是可能只是可读写，非零拷贝中转。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">ssize_t</span> splice(<span class="keyword">int</span> fd_in, <span class="keyword">off_t</span> *_Nullable off_in,</span><br><span class="line">              <span class="keyword">int</span> fd_out, <span class="keyword">off_t</span> *_Nullable off_out,</span><br><span class="line">              <span class="keyword">size_t</span> size, <span class="keyword">unsigned</span> <span class="keyword">int</span> flags);</span><br></pre></td></tr></table></figure><p>Flag 如下：</p><ul><li>SPLICE_F_MOVE<br>  Attempt to move pages instead of copying. 这里的 move 指的是内核页缓存中的物理页面的引用在 fd 之间进行转移。而不需要读出、复制到 user space、写入这样的流程了。<br>  注意，这个 flag 只是一个 hint。如果内核无法移动，则还是需要复制。如果 pipe buffer 不指向整个页面。<br>  The initial implementation of this flag was buggy: therefore starting in Linux 2.6.21 it is a no-op (but is still permitted in a splice() call); in the future, a correct implementation may be restored.</li><li>SPLICE_F_NONBLOCK<br>  Do not block on I/O. This makes the splice pipe operations nonblocking, but splice() may nevertheless block because the file descriptors that are spliced to/from may block (unless they have the O_NONBLOCK flag set).</li><li>SPLICE_F_MORE<br>  More data will be coming in a subsequent splice. This is a helpful hint when the fd_out refers to a socket (see also<br>  the description of MSG_MORE in send(2), and the description of TCP_CORK in tcp(7)).</li><li>SPLICE_F_GIFT<br>  Unused for splice(); see vmsplice(2).</li></ul><h1 id="vmsplice"><a href="#vmsplice" class="headerlink" title="vmsplice"></a>vmsplice</h1><p>splice 主要是服务内核空间中的数据传输，原因是指定的都是 fd 或者 pipe，并不包含用户空间中内存的信息。<br>而 vmsplice 主要服务用户空间和管道之间的数据读写，它们都能实现零拷贝。</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line"><span class="meta">#<span class="meta-keyword">define</span> _GNU_SOURCE         <span class="comment">/* See feature_test_macros(7) */</span></span></span><br><span class="line"><span class="meta">#<span class="meta-keyword">include</span> <span class="meta-string">&lt;fcntl.h&gt;</span></span></span><br><span class="line"></span><br><span class="line"><span class="keyword">ssize_t</span> vmsplice(<span class="keyword">int</span> fd, <span class="keyword">const</span> struct iovec *iov,</span><br><span class="line">                <span class="keyword">size_t</span> nr_segs, <span class="keyword">unsigned</span> <span class="keyword">int</span> flags);</span><br></pre></td></tr></table></figure><p><code>iov</code> 是一个长度为 <code>nr_segs</code> 的数组，表示用户内存中的多段可能不连续的 buffer。</p><p>参数：</p><ul><li><p>SPLICE_F_MOVE<br>  Unused for vmsplice(); see splice(2).</p></li><li><p>SPLICE_F_NONBLOCK<br>  Do not block on I/O; see splice(2) for further details.</p></li><li><p>SPLICE_F_MORE<br>  Currently has no effect for vmsplice(), but may be implemented in the future; see splice(2).</p></li><li><p>SPLICE_F_GIFT<br>  The user pages are a gift to the kernel.<br>  表示用户程序不会修改这段 buffer，否则，page cache 和磁盘中的数据就可能不一致。<br>  将 pages gifting 给内核意味着后面的 splice SPLICE_F_MOVE 能够成功移动 pages。如果不指定，则后续的 splice SPLICE_F_MOVE 必须复制。<br>  数据必须要 page aligned。我理解这里指的是：</p><ul><li><code>iovec[i].iov_base</code> 需要对齐到页</li><li><code>iov_len</code> 需要是页大小的整数倍</li></ul><p>  如果不满足，则退化到 copy 的行为。</p></li></ul><h1 id="Reference"><a href="#Reference" class="headerlink" title="Reference"></a>Reference</h1><ul><li><a href="https://man7.org/linux/man-pages/man2/splice.2.html" target="_blank" rel="noopener">https://man7.org/linux/man-pages/man2/splice.2.html</a></li><li><a href="https://man7.org/linux/man-pages/man2/pipe.2.html" target="_blank" rel="noopener">https://man7.org/linux/man-pages/man2/pipe.2.html</a></li></ul>]]></content>
    
    
    <summary type="html">&lt;p&gt;介绍 Linux 中的零拷贝技术。从 &lt;a href=&quot;/2025/03/09/learn-fuse/&quot;&gt;Fuse 学习&lt;/a&gt; 中独立出来。&lt;/p&gt;</summary>
    
    
    
    
    <category term="Linux" scheme="http://www.calvinneo.com/tags/Linux/"/>
    
    <category term="FileSystem" scheme="http://www.calvinneo.com/tags/FileSystem/"/>
    
  </entry>
  
  <entry>
    <title>My Experience of Building a Hybrid Rust/C++ Project</title>
    <link href="http://www.calvinneo.com/2025/07/21/start-new-rust-project/"/>
    <id>http://www.calvinneo.com/2025/07/21/start-new-rust-project/</id>
    <published>2025-07-21T15:07:22.000Z</published>
    <updated>2026-05-26T08:49:49.290Z</updated>
    
    <content type="html"><![CDATA[<p>Since April 2025, I have been actively contributing to a new Rust–C++ project. Through this work, I have gained many valuable insights. Although I cannot disclose most project details, there are numerous technical challenges worth discussing.</p><p>One of the most notable aspects of this project is that it has been developed alongside the rapid evolution of AI agents, which led us to encounter many pitfalls when practicing vibe coding.</p><a id="more"></a><h1 id="About-Vide-Coding-The-benefits-and-the-pitfalls"><a href="#About-Vide-Coding-The-benefits-and-the-pitfalls" class="headerlink" title="About Vide Coding: The benefits and the pitfalls"></a>About Vide Coding: The benefits and the pitfalls</h1><h2 id="Pitfalls-of-Vibe-Coding"><a href="#Pitfalls-of-Vibe-Coding" class="headerlink" title="Pitfalls of Vibe Coding"></a>Pitfalls of Vibe Coding</h2><p>In early 2025, at the initial stage of our project, one of our core contributors quickly prototyped a demo using Cursor, covering multiple modules such as the read scheduler, index writer, and meta service.</p><p>Traces of this early implementation can still be found in the following pull requests:</p><ul><li><a href="https://github.com/pingcap-inc/tici/pull/48" target="_blank" rel="noopener">https://github.com/pingcap-inc/tici/pull/48</a></li><li><a href="https://github.com/pingcap-inc/tici/pull/673" target="_blank" rel="noopener">https://github.com/pingcap-inc/tici/pull/673</a></li></ul><h2 id="Make-AI-agent-more-focused"><a href="#Make-AI-agent-more-focused" class="headerlink" title="Make AI agent more focused"></a>Make AI agent more focused</h2><h3 id="Why-focused-attention-matters"><a href="#Why-focused-attention-matters" class="headerlink" title="Why focused attention matters"></a>Why focused attention matters</h3><p>Agents have a limited attention budget. When a prompt blends background, implementation details, and review notes, attention is spread thin and the model drifts. Treating the prompt as code and slicing the work into sub goals keeps the highest-signal spec in focus, which improves determinism, reduces requirement misses, and lowers evaluation cost.</p><h3 id="How-to-draw-attention-of-AI-agent"><a href="#How-to-draw-attention-of-AI-agent" class="headerlink" title="How to draw attention of AI agent"></a>How to draw attention of AI agent</h3><p>We can treat prompt as code. This is not a new idea in the industry, but many teams apply it unevenly. A common practice is to store prompts as markdown or templates under version control, so they can be reviewed, diffed, and rolled back like any other artifact. Some teams go further and build a “prompt registry” or config service to version prompts outside the codebase, and pair it with evaluation suites that act like unit tests for prompts (golden outputs, A/B runs, regression checks). Others embed prompts directly in application code as constants, which makes deployment easy but tends to hide intent and lose reviewability. The shared direction is clear: treat prompts as first-class assets with explicit structure, reviews, and tests.</p><p>I wrote a PR, <a href="https://github.com/pingcap-inc/tici/pull/692/files" target="_blank" rel="noopener">https://github.com/pingcap-inc/tici/pull/692/files</a>, where the core file(I name it spec), is <code>prompts/0001-gc-cdc.md</code>. That file is the prompt itself, and it is committed with the PR, so anyone can start a session by loading the same prompt. It becomes <strong>versioned</strong>, <strong>diffable</strong>, and <strong>reviewable</strong> like real code, and the team no longer depends on a hidden chat history. Also, we can divide the whole goal into several sub goals and let AI agents to implement different sub goals sequentially or in parallel. And we don’t need to tell the AI the context everytime.</p><p>I think this can effectively make the AI agent more focused on what it needs to do, so as to generate bettwer codes with less resources.</p><p><strong>Reviewers</strong> can write feedback directly in the <code>Reviews</code> section of the prompt doc, and I can pull that back into the next round of vibe. This workflow makes context length much less of a concern because the prompt is the canonical spec. A minimal shape looks like this:</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br></pre></td><td class="code"><pre><span class="line">// The content of a spec file</span><br><span class="line"># Goal</span><br><span class="line"></span><br><span class="line">...</span><br><span class="line"></span><br><span class="line">## Sub goal 1</span><br><span class="line"></span><br><span class="line">In this sub goal, you need to ...</span><br><span class="line"></span><br><span class="line">### Programming Style</span><br><span class="line"></span><br><span class="line">### Musts</span><br><span class="line"></span><br><span class="line">### Tests</span><br><span class="line"></span><br><span class="line">## Reviews</span><br></pre></td></tr></table></figure><p>This spec file can be compacted automatically during refactors.</p><h2 id="Rust-as-the-language-of-“Vide-Coding-Era”"><a href="#Rust-as-the-language-of-“Vide-Coding-Era”" class="headerlink" title="Rust as the language of “Vide Coding Era”?"></a>Rust as the language of “Vide Coding Era”?</h2><p>As Rust becomes increasingly adopted in the “vibe coding era”, it does offer stronger guarantees against concurrency and memory errors. However, it is still too early to say that Rust is THE ONE.</p><p>Such an AT-Agent native language will likely consist of at least three distinct sub-languages:</p><ul><li>One for expressing intent<br>  This is the most critical language, because <strong>Developers</strong> need a more efficient way to understand what AI agents have actually done.</li><li>One for validating correctness, or so called DoD definition<br>  The language corresponds to the intent language, and is designed to describe test workflows more efficiently. <strong>Developers</strong> and <strong>Reviwers</strong> can fine-tune this part of the code to guide AI agents to generate correct code for corner cases.</li><li>One for concrete implementation<br>  This language is responsible for concrete implementation. Many AI agents, such as Codex, can already handle this layer well, as they rarely make trivial mistakes. <strong>Developers</strong> do not need to review this part of the code frequently, as other AI agents can handle the review instead.</li></ul><p>Only with this separation can we truly balance readability, robustness, and long-term maintainability. Furthermore, many of the current libraries can be rewritten to be more friendly to AI agent.</p><h3 id="The-goal-command"><a href="#The-goal-command" class="headerlink" title="The /goal command"></a>The /goal command</h3><p>此后，大多数 agent 都提供了 /goal 这个功能，其目的也是为了强化目标驱动；</p><figure class="highlight plain"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br></pre></td><td class="code"><pre><span class="line">human writes steps</span><br><span class="line"></span><br><span class="line">==&gt;</span><br><span class="line"></span><br><span class="line">human specifies outcome</span><br><span class="line">agent searches for path</span><br></pre></td></tr></table></figure><p>不过，/goal 将我们需要的语言展示得更为具体了。</p><h2 id="Use-Skills"><a href="#Use-Skills" class="headerlink" title="Use Skills"></a>Use Skills</h2><p>I introduced several SKILLs into our new project. For example, <a href="https://gist.github.com/CalvinNeo/76811a9fbdd58d1bd271f17004051160" target="_blank" rel="noopener">ManualTest</a> enables AI agents to execute manual test cases automatically.</p><h1 id="Implement-tests"><a href="#Implement-tests" class="headerlink" title="Implement tests"></a>Implement tests</h1><h2 id="Hierarchy-of-tests"><a href="#Hierarchy-of-tests" class="headerlink" title="Hierarchy of tests"></a>Hierarchy of tests</h2><h3 id="The-problem"><a href="#The-problem" class="headerlink" title="The problem"></a>The problem</h3><p>Because each TiDB component is maintained in a separate repository, breaking changes in one component require coordinated adaptations across multiple repositories. Unfortunately, such adaptations cannot be performed atomically and are often non-trivial. While compilation flags or configuration options can sometimes be used to temporarily disable new features, this strategy is not always applicable. In particular, interface changes such as FFI definitions may break compatibility immediately. </p><p>In our project, an end-to-end (e2e) test starts a full cluster and asserts it from a client’s perspective, which in our case means sending SQL queries to the database service. Some of these tests are included in our CI pipeline. However, CI-based e2e tests cannot reliably detect adaptation issues. This creates a classic catch-22: resolving an adaptation problem requires updating all related components, yet the e2e tests cannot pass while you are still fixing the first component. As a result, most e2e tests are deferred to what we call the “daily tests”.</p><p>Nevertheless, we still need a subset of e2e tests in the CI pipeline. Although these tests may occasionally produce false positives due to compilation or adaptation issues, they provide valuable systematic checks to ensure that a new commit in one component does not break existing rules or behaviors. Deferring such checks to daily tests would be disastrous, as it makes bugs significantly harder to triage. When issues accumulate over time, the project can easily fall into a “bug jail,” where fixing new problems becomes increasingly expensive.</p><p>It is also worth noting that integration tests cannot practically detect all logical bugs. In many cases, module owners write integration tests mainly to verify that their own modules work with others, while overlooking the impact their changes may have on the system as a whole. This issue becomes even more critical when AI agents are used to refactor our code, as we need safeguards to ensure that unexpected behavior does not compromise the foundation of the project.</p><p>During the development and PoC stage of our project, several critical issues occurred because the tests were not correctly implemented, including:</p><ul><li>Module A uses the API of Module B in a wrong way. Neither is there integration test of Module A, nor its scene is covered in the e2e test.</li><li>Module C fails to verify a corner case, which is later caught by my <strong>embedded e2e test</strong> (introduced later). This kind of error is easy to be ignored, because it passes all tests except one. However, that single test protects our system from an availability failure caused by a deadlock in Module C.</li><li>Another component changed its convention for constructing a field in an RPC request without informing us, which caused the system to malfunction at the SQL layer and made the issue difficult to investigate. This problem was also detected by my <strong>embedded e2e test</strong>.</li></ul><h3 id="The-layers-of-tests"><a href="#The-layers-of-tests" class="headerlink" title="The layers of tests"></a>The layers of tests</h3><p>We can organize the tests for Component A into several layers. Each layer targets specific categories of bugs, and issues detected at lower layers should not propagate to higher layers.</p><ul><li>Daily regression: full system tests for cross-component compatibility only. However, no bugs originating from Component A itself should reach this level. We must proactively investigate daily test failures to avoid falling into a bug jail.</li><li>E2e tests with real components: We may occasionally allow skipping tests at this layer, because cross-component checks can fail due to upstream changes, as discussed earlier. However, bugs that originate within Component A itself must not propagate to this level.</li><li>Integration tests:<ul><li>Embedded e2e: at the component boundary, using mocked RPC/status/FFI calls; required when interface semantics change; validates external behavior and isolates compatibility issues.</li><li>Module integration: per-module behavior with or without mocks; required for new features or refactors; may need test framework enhancements.</li></ul></li><li>Unit tests: unit or local integration tests within a single module; most of these can be handled by AI agents.</li></ul><h3 id="The-embeded-e2e-test"><a href="#The-embeded-e2e-test" class="headerlink" title="The embeded e2e test"></a>The embeded e2e test</h3><p>This idea is based on the observation that a component’s behavior is defined by how it communicates with other components, through RPC, FFI, shared memory, and similar mechanisms.</p><p>Therefore, mocking these communications in integration tests provides the following benefits:</p><ul><li>We don’t need to start a full cluster, so we won’t face the adaption problem.</li><li>If an adaptation issue occurs, it can be easily reproduced at this level. This not only simplifies the debugging process, but also increases our confidence in the code.</li><li>This test treats our program as a black box, which makes it easier to implement because we do not need to understand how each module is implemented. These tests are expected to remain stable unless the interfaces or communication frameworks change.</li></ul><h3 id="Tests-as-the-Backbone-of-Vibe-Coding"><a href="#Tests-as-the-Backbone-of-Vibe-Coding" class="headerlink" title="Tests as the Backbone of Vibe Coding"></a>Tests as the Backbone of Vibe Coding</h3><p>In a Vibe Coding workflow, tests become the primary communication channel between intention and code. Among all types of tests, the embedded end-to-end (e2e) tests play a more and more important role.</p><p>Unlike unit tests, which specify local behavior, or integration tests, which usually verify a limited subsystem, my embedded e2e tests define system-level behavioral contracts. They describe what the system should do rather than how it should do it. This makes them naturally aligned with Test-Driven Development (TDD): they serve as executable specifications that drive the implementation.</p><h1 id="Systematic-choices"><a href="#Systematic-choices" class="headerlink" title="Systematic choices"></a>Systematic choices</h1><h2 id="Thread-or-coroutine"><a href="#Thread-or-coroutine" class="headerlink" title="Thread or coroutine?"></a>Thread or coroutine?</h2><p>Benefits of using tokio:</p><ul><li>Smaller memory cost, so we can create more coroutines.</li><li>Context switch is faster because there is no syscall.</li></ul><p>Pitfalls of using tokio:</p><ul><li>We cannot control the scheduling strategy of tokio’s runtime. For example, we cannot assign a priority to a specific task, nor can we limit the CPU quota of a particular class of tasks.</li><li>Switching to async code is often painful, as even the simplest function may become suspendable due to the use of <code>tokio::sync</code> locks.</li><li>It is hard to investigate deadlock / starvation problems.</li><li>Hard to use itertools. <code>futures::stream</code> can help, but it generates complex types.</li></ul><h3 id="Use-seperated-Runtime-for-different-task-pool"><a href="#Use-seperated-Runtime-for-different-task-pool" class="headerlink" title="Use seperated Runtime for different task pool?"></a>Use seperated Runtime for different task pool?</h3><p><code>Runtime</code> can only be created outside the “async context” of tokio. So if we need to use tuned <code>Runtime</code>s, we have to create them in advance. This involves lots of refactors.</p><h3 id="Propagate-the-panic-outward"><a href="#Propagate-the-panic-outward" class="headerlink" title="Propagate the panic outward"></a>Propagate the panic outward</h3><p>We must pay attention to panics inside the actor’s message loop: the handler, whether a thread or a coroutine, will only surface the panic when it is eventually joined, by which time the failure may have gone unnoticed for too long. What I recommend is to:</p><ul><li><p>Employ the <code>panic_hook</code> to capture the exact scene where things go wrong.</p>  <figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br></pre></td><td class="code"><pre><span class="line">panic::set_hook(<span class="built_in">Box</span>::new(|info| &#123;</span><br><span class="line">    eprintln!(<span class="string">"Task panicked: &#123;&#125;"</span>, info);</span><br><span class="line">    <span class="built_in">println!</span>(<span class="string">"Task panicked: &#123;&#125;"</span>, info);</span><br><span class="line">&#125;));</span><br></pre></td></tr></table></figure></li><li><p>Eliminate <code>unwrap</code>s and <code>expect</code>s</p>  <figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br></pre></td><td class="code"><pre><span class="line"><span class="meta">#![cfg_attr(not(test), deny(clippy::unwrap_used))]</span></span><br><span class="line"><span class="meta">#![cfg_attr(not(test), deny(clippy::expect_used))]</span></span><br></pre></td></tr></table></figure></li></ul><h2 id="Shared-Memory-or-Actor-model"><a href="#Shared-Memory-or-Actor-model" class="headerlink" title="Shared Memory or Actor model?"></a>Shared Memory or Actor model?</h2><p>If we use the coroutine runtime, we may need to decide how to handle race conditions.</p><h3 id="Why-are-“deadlock”s-so-hard-to-diagnose-when-using-coroutines"><a href="#Why-are-“deadlock”s-so-hard-to-diagnose-when-using-coroutines" class="headerlink" title="Why are “deadlock”s so hard to diagnose when using coroutines?"></a>Why are “deadlock”s so hard to diagnose when using coroutines?</h3><ol><li>There is neither a wait-for graph in the coroutine runtime nor one in the OS<br> <code>await</code> does not block a thread, so we can’t find anything with gdb/strace/perf.<br> Meanwhile, these “deadlocks” are hard to be detected because they appears that there is no CPU, no blocking thread, and the program is in a “vegetative state”.<br> Coroutine frameworks like <code>tokio</code> provides some o11y tools, however, they are hard to use, and have performance overhead.</li><li>No actual “deadlock”<br> These stalls are mostly “waiting for a train at a bus stop” errors. For example, we may read from a channel which will never be written, which is an easy mistake when we bail on an error without calling <code>.send()</code> first.<br> So we recommend to send a <code>Result&lt;T&gt;</code>, and implement a <code>Drop</code> trait that automatically sends <code>Err(Error::DropWithoutReport)</code> as a last-minute remedy.</li><li>No actual “stack”<br> Coroutines don’t carry a real stack. When they hit an await they yield a continuation, and that continuation may be resumed on the same or a different thread.</li></ol><h3 id="tokio-RwLock-or-std-sync-Mutex"><a href="#tokio-RwLock-or-std-sync-Mutex" class="headerlink" title="tokio::RwLock or std::sync::Mutex?"></a>tokio::RwLock or std::sync::Mutex?</h3><p>There is a public belief that we have to always use tokio locks in asynchronous code. However, according to the <a href="https://docs.rs/tokio/latest/tokio/sync/struct.Mutex.html#which-kind-of-mutex-should-you-use" target="_blank" rel="noopener">reference</a> of tokio, it is ok and better to use synchronous locks such as <code>std::sync::Mutex</code> or <code>parking_lot::Mutex</code>.</p><p>I’d like to refer to these cases as “atomic access structures”, because they all follow the following patten:</p><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br></pre></td><td class="code"><pre><span class="line"><span class="class"><span class="keyword">struct</span> <span class="title">Wrapped</span></span> &#123;</span><br><span class="line">    inner: Mutex&lt;<span class="built_in">String</span>&gt;,</span><br><span class="line">&#125;</span><br><span class="line"></span><br><span class="line"><span class="keyword">impl</span> Wrapped &#123;</span><br><span class="line">    <span class="keyword">pub</span> <span class="function"><span class="keyword">fn</span> <span class="title">change_inner</span></span>(&amp;<span class="keyword">self</span>, s: <span class="built_in">String</span>) &#123;</span><br><span class="line">        <span class="keyword">self</span>.inner.lock().expect(<span class="string">""</span>) = s;</span><br><span class="line">    &#125;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>The key point of this code is to avoid directly exposing the lock itself: we must not allow external callers to access it, and we must atomically release the lock after mutating the protected value. The underlying rationale is that we must not allow a coroutine to “sleep” while holding the lock, as this guarantees that no deadlocks will occur, because:</p><ul><li>If a coroutine holds the lock, it will not “sleep”, because the code <code>change_inner</code> is structured to avoid calling <code>.await</code> while the lock is held. Moreover, the executor thread will not sleep either, since it is not waiting on any condition.</li><li>If a coroutine does not hold the lock, it can eventually acquire it, because the current holder will release the lock promptly. And of course, the lock is released before any suspension point.</li></ul><p>For synchronous lock:</p><ul><li><code>std::sync::Mutex</code> supports poisoning. If a thread panics while holding the lock, future <code>lock()</code> calls return a <code>PoisonError</code>, forcing the caller to acknowledge that the protected state may be inconsistent. However, in most cases such an error can’t be processed, it will eventually lead to a panic, which is not elegant. There are also some other choices, simply accepting the risk is similar with <code>parking_lot::Mutex</code>, and resetting the state is not hazard if there are other threads waiting for this lock.</li><li><code>parking_lot::Mutex</code> is faster and smaller in many workloads (especially uncontended or lightly contended), but it does not poison. This often turns a hard failure into a latent, harder-to-debug corruption.</li></ul><p>The lack of poisoning is exactly why <code>parking_lot::Mutex</code> can be unsafe at the <em>logic</em> level. If a panic happens after partially mutating the protected value, the invariant is already broken. So, unless you can guarantee panic-free critical sections or have a clear recovery path, prefer the standard mutex to make invariant breaks visible. I think the best practice is to abort if any thread panics, which can be easily done with panic hooks.</p><h3 id="Implementing-an-“incomplete”-actor-mode"><a href="#Implementing-an-“incomplete”-actor-mode" class="headerlink" title="Implementing an “incomplete” actor mode"></a>Implementing an “incomplete” actor mode</h3><p>In the traditional actor model, each actor node encapsulates its own private data. However, this model is difficult to implement because:</p><ul><li>To rebalance data across nodes, we must introduce new message types and corresponding handlers.</li><li>Inspecting the internal state of actor nodes is difficult.</li></ul><p>So, as a simpler alternative, we can:</p><ul><li>Use a concurrent hash map to store all data, with each actor node mutating a portion of the map.</li><li>Allow other components to read or inspect entries in the concurrent hash map. Such inspectors can not mutate the entries, and their accessment must be atomic.</li></ul><p>A preferred candidate for the hash map is <code>DashMap</code>. Although this structure frees us from requiring <code>&amp;mut self,</code> most of its methods return a <code>Ref</code> or <code>RefMut</code> that holds a lock guard, so incorrect usage can lead to deadlocks. The following codes shows a simple example.</p><figure class="highlight rust"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br></pre></td><td class="code"><pre><span class="line"><span class="meta">#[test]</span></span><br><span class="line"><span class="function"><span class="keyword">fn</span> <span class="title">test_dashmap</span></span>() &#123;</span><br><span class="line">    <span class="keyword">let</span> map = DashMap::new();</span><br><span class="line">    map.insert(<span class="number">1</span>, <span class="number">1</span>);</span><br><span class="line">    map.insert(<span class="number">2</span>, <span class="number">2</span>);</span><br><span class="line"></span><br><span class="line">    <span class="keyword">for</span> entry <span class="keyword">in</span> map.iter() &#123;</span><br><span class="line">        <span class="built_in">println!</span>(<span class="string">"&#123;&#125; -&gt; &#123;&#125;"</span>, entry.key(), entry.value());</span><br><span class="line">        map.insert(<span class="number">3</span>, <span class="number">3</span>);</span><br><span class="line">    &#125;</span><br><span class="line"></span><br><span class="line">    <span class="built_in">println!</span>(<span class="string">"test end"</span>);</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>There is a simple yet effective way to detect potential issues in our code: use <code>#[tokio::test]</code> instead of <code>#[tokio::test(flavor = &quot;multi_thread&quot;)]</code>. With the single-threaded runtime, the program will fail immediately if a coroutine “sleeps” while holding a lock.</p><h2 id="The-linking-problem"><a href="#The-linking-problem" class="headerlink" title="The linking problem"></a>The linking problem</h2><h3 id="FFI"><a href="#FFI" class="headerlink" title="FFI"></a>FFI</h3><p>TODO</p><h3 id="How-to-support-TLS"><a href="#How-to-support-TLS" class="headerlink" title="How to support TLS?"></a>How to support TLS?</h3><p>TODO</p><h2 id="Online-config-change"><a href="#Online-config-change" class="headerlink" title="Online config change"></a>Online config change</h2><p>There are some ways to update configs without restarting the program:</p><ul><li>For every actor, introduce a new <code>UpdateConfig</code> event, and handle it in the message loop.</li><li>Using <code>arc_swap</code>.</li></ul><p>I don’t think the service itself should persist the updated configuration to the config file. Instead, this should be handled by the operator.</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;Since April 2025, I have been actively contributing to a new Rust–C++ project. Through this work, I have gained many valuable insights. Although I cannot disclose most project details, there are numerous technical challenges worth discussing.&lt;/p&gt;
&lt;p&gt;One of the most notable aspects of this project is that it has been developed alongside the rapid evolution of AI agents, which led us to encounter many pitfalls when practicing vibe coding.&lt;/p&gt;</summary>
    
    
    
    
    <category term="C++" scheme="http://www.calvinneo.com/tags/C/"/>
    
    <category term="arch" scheme="http://www.calvinneo.com/tags/arch/"/>
    
    <category term="Rust" scheme="http://www.calvinneo.com/tags/Rust/"/>
    
    <category term="Articles" scheme="http://www.calvinneo.com/tags/Articles/"/>
    
    <category term="agent" scheme="http://www.calvinneo.com/tags/agent/"/>
    
  </entry>
  
  <entry>
    <title>乒乓球训练纪实</title>
    <link href="http://www.calvinneo.com/2025/06/11/table-tennis/"/>
    <id>http://www.calvinneo.com/2025/06/11/table-tennis/</id>
    <published>2025-06-11T15:07:22.000Z</published>
    <updated>2026-05-17T18:15:57.629Z</updated>
    
    <content type="html"><![CDATA[<p>因为五一节打羽毛球把膝盖打出问题了，现在主要学习乒乓球了</p><a id="more"></a><h1 id="Day-1"><a href="#Day-1" class="headerlink" title="Day 1"></a>Day 1</h1><p>主要修正了下反手动作。</p><p>这里，我最重要的是不要架肘。打的时候，可以用左手稍微按着一点左边的大臂。<br>然后，乒乓球反手是小臂带动大臂，大臂基本上不用特别发力，而是让小臂画 1/4 的圆弧。<br>在打完之后，需要还原。击球时，击球点不要太到台内，仿佛把整个手都要伸出去一般。<br>击球的是高点期。注意，不要急，而是等球到位之后再打，实际上质量更好。<br>当球落点朝着左边或者右边偏斜的时候，可以考虑靠只移动上半身来接。<br>作为初学者，不需要手腕有特别大的扭转。</p><p>-&gt; 初学者，可以采用下蹲马步，然后再站起来这样。但是最终，是要一直处于一种扎马步的状态，主要是前脚掌着地。然后用一些垫步还是啥的完成重心转换。</p><h1 id="Day-2"><a href="#Day-2" class="headerlink" title="Day 2"></a>Day 2</h1><p>主要纠正了下正手动作。</p><p>我正手有几点问题：</p><ul><li>撅手腕，教练说撅着手腕可能是因为击打点太靠前，以至于要“等球：</li><li>大臂架着</li><li>大臂还喜欢架着往后拉，实际上应该是转腰，而不是动大臂</li><li>握拍应该虎口直接对着拍子的“侧棱”，类似于羽毛球一样，我可能会喜欢转一点</li></ul><h1 id="Day-3"><a href="#Day-3" class="headerlink" title="Day 3"></a>Day 3</h1><p>继续是正手和反手的练习，下雨，只有 1.5h。</p><p>提到了一些细节：</p><ul><li>反手可以尝试蹲起，然后站立这种击打方式</li></ul><p>录制了一些视频，视频中可以发现，还是喜欢撅着手腕。</p><h1 id="Day-4"><a href="#Day-4" class="headerlink" title="Day 4"></a>Day 4</h1><p>继续是正手和反手的练习，有事，只有 1.5h。</p><h1 id="Aug-21"><a href="#Aug-21" class="headerlink" title="Aug 21"></a>Aug 21</h1><p>纠正拉手问题：</p><ul><li>拉手是肘往后牵引，这个要注意避免。如果拉手，那么正手打出的球可能带有侧旋</li><li>可以考虑自己夹着一个矿泉水瓶来打正手，能减少拉手问题。</li></ul><h1 id="Sep-4"><a href="#Sep-4" class="headerlink" title="Sep 4"></a>Sep 4</h1><p>强调转胯的作用：</p><ul><li>是那种压着腹股沟的感觉</li><li>要通过转胯而不是拉手来准备正手的引拍</li></ul><h1 id="Sep-11"><a href="#Sep-11" class="headerlink" title="Sep 11"></a>Sep 11</h1><p>首先询问了还原和启动的关系。教练的意思是攻球，挥拍到前面之后，可以先不动。等对面球出来之后，再判断，然后直接启动。比如如果是正手球，就可以直接侧身准备击球。还原和启动是一个动作完成了。</p><p>然后是关于撞和磨的关系的讨论，教练认为还是要先撞后磨的。</p><p>这次练了下，感觉最重要的还是放松。</p><h1 id="Sep-19"><a href="#Sep-19" class="headerlink" title="Sep 19"></a>Sep 19</h1><p>因为重新看了下医生，说暂时不要运动，因此这次就练了下搓球。</p><p>搓球的话主要几点：</p><ul><li>腿要伸进去</li><li>重心一定要压低压低！可以感受到脸要贴着桌子的感觉</li><li>击球点要低，基本上是二跳要落下，就要完蛋的那个时候搓。但尽管如此，人是要先上去的，所以这里有一个停顿</li></ul><h1 id="Sep-26"><a href="#Sep-26" class="headerlink" title="Sep 26"></a>Sep 26</h1><p>本周停止</p><h1 id="Oct-22"><a href="#Oct-22" class="headerlink" title="Oct 22"></a>Oct 22</h1><p>搓球：</p><ul><li>脚上，身体也要上，要定住，然后在搓球</li><li>我搓球的出球有的太高，有的下网了。这个并不是击球点不对，击球点始终是要在下降期的（对于我的水平而言）。搓球的出球高地，主要是要判断来球的旋转，来确定板型。</li><li>劈长是低点，摆短要抢高点。fzh 跟我说劈长也要高点，后面教练跟我说这也不错，主要是为了出球的一致性，毕竟劈长是一个进攻的技术，如果对面起球的质量不高，那么就能一板打死，高点劈长的效果就体现出来了。但是这个比较难，我暂时不需要掌握。</li><li>无论如何，搓球球拍要在下面。搓球也是要往前，去点的那种感觉的，不是往下砍。</li></ul><p>攻球：</p><ul><li>正手攻球，其实只要侧身就行。我不要想着要侧很多，因为这样真的容易侧 180 度，类似于羽毛球那样。或者，就是直接拉手了。我理解这里跟羽毛球里面刘辉说的一样，就是我们对于侧身的感知太不敏感了，往往实际侧了很多，我们只感觉到侧了一点。</li></ul><h1 id="Oct-30"><a href="#Oct-30" class="headerlink" title="Oct 30"></a>Oct 30</h1><ul><li>握拍得贴着虎口。因为虎口很大，所以是得靠大拇指一侧，而不是食指一侧</li><li>搓球的时候，中心前压，不是说要驼背，而是腹部发力</li><li>打球的时候，可以尝试只用前脚掌，把重心放到前脚掌上</li><li>摆速的时候，手不要拉</li><li>搓球的时候，中心要低，但不要趴桌上</li></ul><h1 id="Nov-13"><a href="#Nov-13" class="headerlink" title="Nov 13"></a>Nov 13</h1><p>正手虽然动作大，但是击球点靠后，所以空中时间长。</p><p>搓球亮板子是为了接触面积大。</p><h1 id="Nov-21"><a href="#Nov-21" class="headerlink" title="Nov 21"></a>Nov 21</h1><ul><li>搓球的时候，不要靠大臂往前戳。而是要以肘为轴心，小臂往前，大臂是不需要往前伸的</li></ul><h1 id="Nov-27"><a href="#Nov-27" class="headerlink" title="Nov 27"></a>Nov 27</h1><p>提了几点：</p><ul><li>反手搓球，手也要拱着</li></ul><h1 id="Dec-4"><a href="#Dec-4" class="headerlink" title="Dec 4"></a>Dec 4</h1><ul><li><p>反手搓球，直接上右脚就行，没必要往左边跨一步</p></li><li><p>反手搓球和正手搓球，手都要拱着</p></li><li><p>搓球的时候，拍子不要很立，或者说不要翘起来</p></li><li><p>正手搓球的时候，手可以送一点出去。我理解就是搓球的时候不要像触电一样</p></li><li><p>反手拨球，手也不要翘起来</p></li></ul><h1 id="Dec-11"><a href="#Dec-11" class="headerlink" title="Dec 11"></a>Dec 11</h1><p>正手攻球打球有几点：</p><ul><li>上旋球，如果弧线比较低，也不需要提，直接打过去就行</li><li>攻球的时候，不仅要注意手不要翘，不要拉手。手腕最好跟搓球一样，能拱一点起来。整体来讲，就是要接触面积大。</li></ul><p>反手攻球，发现有时候会打下网，原因还是我手腕没有偏拱，而是有点把手翘起来的感觉。我在微信里面备注了个图，可以参考下。</p><p>另外，进一步看了正手搓球：</p><ul><li>重心还是要低，这里强调了一下，肘部也要低。我打球的时候，可能是提着肘，然后手在下面。其实应该肘和手都在下面。</li><li>可以想象，准备姿势，重心压低。这个时候，手臂是很贴近台面的。所以，正手搓球的时候，只需要把手臂打横就行了。</li><li>搓球也需要往前送，正手搓球不要往自己身体收前臂，而是往前送。</li><li>其实搓球搓起来，球不太容易很贴网的。</li></ul><h1 id="Jan-7"><a href="#Jan-7" class="headerlink" title="Jan 7"></a>Jan 7</h1><p>中间放假缺了一节课，另外和 fzh 又打了几场球，感觉正手搓球有点差了。主要体现在：</p><ul><li>伸手太靠近身体</li><li>身体的发力太多了，手上的动作少了</li><li>小臂不要翘上去，搞得像一个 V 字一样，还是要放下来一点</li></ul><p>在攻球方面，感觉膝盖好了不少。</p><p>然后初步学习一下反手拉球，感觉我自己的问题是：</p><ul><li>手肘架太高了，更像拧</li><li>可能是怕刮到球台，自己退台太多了</li></ul><p>这里的一个重点在于：</p><ul><li>手腕一定要引拍，让球拍对着自己</li><li>拉球完毕后，手腕应该恢复类似正常的攻球状态，不要往上翘。其实往上翘也是我之前攻球常见的一个错误问题</li></ul><p>我理解反手拉球随着你引拍是在身体左边、中间还是右边，动作的大小和幅度不太一样，但是框架是一样的，就是手腕要制造摩擦。</p><p>Jan 8 的时候又和 fzh 语音交流了一下。他的意思是反手拉球其实最重要的是能拉到别人搓过来的比较快并且贴身体的球。基本上退台是不会反手拉的。</p><h1 id="Jan-30"><a href="#Jan-30" class="headerlink" title="Jan 30"></a>Jan 30</h1><p>今天又是只有我一个人。</p><p>搓球：</p><ul><li>正手搓球不要在身体前面，而是要打开手臂，在身体右前方击球，原因是因为要有一致性。这样，站在那个位置，处理方式会有很多，不会给人感觉你就一定是搓球。不过打开手臂之后，还是要往前的，不能往身体方向收手臂。这个刻意注意一下就行。</li><li>另一种注意的办法是，搓球的时候可以有个停顿。这里的停顿不是说你就把拍子放在那里等球来了，我们肯定是要向着球制造摩擦的。但是就和羽毛球那样，得有一种“定”的感觉。或者可以认为伸手臂和往前搓实际上可以分解为两个动作。</li></ul><p>拉球：</p><ul><li>反手拉球直接跨步调整位置，不需要滑步。这是因为反手的地方就那么一块，滑步的动作太大了。</li></ul><p>另外，我觉得拉球我的问题主要是：</p><ul><li>不要边动边打球。脚步站好的时候，最好重心也调整好，比如该蹲那时候就蹲了，不要一边挥拍一边往下蹲，这样很难瞄准。</li><li>拍子要在下面，从下往上挥动。</li><li>拍子不要往上走，往上是通过手腕那个旋转来做的。手臂就是放松，自然展开就行。</li></ul><h1 id="Feb-12"><a href="#Feb-12" class="headerlink" title="Feb 12"></a>Feb 12</h1><p>学了下劈长。感觉这个很讲究瞬间发力：</p><ul><li>如果小臂移动太多，那么就容易出界。</li><li>如果手腕移动太多，比如最后往上翘了，那么也容易出界。</li></ul><p>另外今天体验了一下摆短。我确实可以摆起来，但是是卸力把球托过去的，并没有搓很转，所以没有什么威胁。</p><p>今天尝试了下拉球，就是彻底把手放下来，然后就发现有个明显的好处，就是如果我手自然下垂，那么我就没必要很刻意地去内旋我的手腕了，反而更轻松。</p><h1 id="Mar-19"><a href="#Mar-19" class="headerlink" title="Mar 19"></a>Mar 19</h1><p>继续是反手拉球：</p><ul><li>拉球是小要外旋，但是是往前，而不是要往上。</li><li>最重要的是，下降期拉球，如果抢了，就上不了。</li></ul><p>继续是搓球：</p><ul><li>我发现如果我每次搓球动作好的时候，都会碰到桌子。教练说没问题，反正我甚至可以就这样找感觉。</li><li>另外搓球就是手要打开。无论是正反手。手和胳膊都不要太硬。</li><li>正手搓球还有一点是，手不要缩着，搓球的点，应该更靠外一点。</li></ul><h1 id="Mar-26"><a href="#Mar-26" class="headerlink" title="Mar 26"></a>Mar 26</h1><h1 id="Apr-1"><a href="#Apr-1" class="headerlink" title="Apr 1"></a>Apr 1</h1><p>正手搓球：</p><ul><li>搓球正手手腕还是要打开。</li><li>然后不要整个手臂都往前捅，只是手腕一个小动作。</li></ul><p>正手拉球：</p><ul><li>不要整个手臂都自然下垂。主要还是手肘要夹着，也不要架着。然后小臂要下垂，然后可以稍微向后引拍。</li></ul><h1 id="Apr-16"><a href="#Apr-16" class="headerlink" title="Apr 16"></a>Apr 16</h1><h1 id="Apr-24"><a href="#Apr-24" class="headerlink" title="Apr 24"></a>Apr 24</h1><p>最近反手好了很多，包括跟 fzh 练的时候也能感受到：</p><ul><li>主要是我现在不会手腕往上翘了，而是往内收，这样摩擦会好点。</li></ul><p>今天和一个高中生多球。感觉上：</p><ul><li>速度起来之后，步伐就不太行。侧身感觉侧不过来，还是手先动。</li><li>搓球很容易冒高，教练的意思是这个球就不转，所以容易搓高。</li><li>搓后接着拉球的时候，步伐吃力。</li></ul><h1 id="May-9"><a href="#May-9" class="headerlink" title="May 9"></a>May 9</h1><ul><li>攻球要打开，也就是说拍子要更往外站开点，可以类比正手搓球。</li><li>攻球可以夹着大臂，拉球不行。</li><li>拉球手腕不要往上翘。</li><li>拉球时候，拍头和小臂平齐就行。</li><li>拉球的时候，类似于用肩膀画弧，但是不要甩大臂。收小臂，把大臂可以带出去。往前不要往左。</li></ul><h1 id="May-13"><a href="#May-13" class="headerlink" title="May 13"></a>May 13</h1><ul><li>正手搓接反手拉。搓球完要回一步。可以理解为搓球的时候右脚就蹬住，搓完了直接用力就能回。然后，看回球的情况，如果短，就可以直接处理，或者特别短的话就上步。如果需要反手拉球，则可以往后垫一步，同时手放下去准备拉球。</li><li>正手拉球，引拍的步伐不要往前走，而是主要体现为要把重心放到右脚上，并且要侧身。</li><li>正手拉球，不要拉完人往后仰，而是重心压在前面。</li><li>正手拉球，前臂要更靠外，可以有快速收小臂的感觉，但是小臂不要收到左边，而是要往前。拉球的整个动作都是要往前的。</li></ul>]]></content>
    
    
    <summary type="html">&lt;p&gt;因为五一节打羽毛球把膝盖打出问题了，现在主要学习乒乓球了&lt;/p&gt;</summary>
    
    
    
    
    <category term="运动" scheme="http://www.calvinneo.com/tags/运动/"/>
    
  </entry>
  
  <entry>
    <title>关西2</title>
    <link href="http://www.calvinneo.com/2025/04/08/kensai-2/"/>
    <id>http://www.calvinneo.com/2025/04/08/kensai-2/</id>
    <published>2025-04-08T12:06:11.000Z</published>
    <updated>2025-04-12T08:20:49.168Z</updated>
    
    <content type="html"><![CDATA[<p>趁着清明节又去了一趟关西。本以为是度假，但实际上累得要死。</p><a id="more"></a><h1 id="D1"><a href="#D1" class="headerlink" title="D1"></a>D1</h1><p>这一次从南京直飞大阪。机票照例没有提前多久定，但是两个人往返也才 4k 出头，相当便宜。宾馆就是贵的离谱了。京都的宾馆单人 1.5k，双人接近 4k 感觉简直在抢钱。我先在大阪定了 2 天 700 左右的。然后又订了一天姬路和大阪的。然后最后一天住哪不清楚，大概先这样。</p><p>关西机场入国变得麻烦多了，这次虽然没有坐小火车（吉祥航空），但入关排队感觉就花了将近一个小时了。在飞机上被发了入境单要填，我觉得挺麻烦的，现在都是电子化了啊。结果到了入境口发现有 abcd 四个 route，但完全不知道区别是什么。一个国际机场居然没有英文的说明，工作人员也只说日语和简单的英语。Anyway，他有个扫描护照的机器，感觉挺方便的。我用机器扫完，然后就再往前排队等人工。走到一半发现大家还有个 QR code 也不知道是啥，但是这些人扫完 QR 之后又要填一遍纸单子。。。</p><p>反正轮到我我就说 I have no QA code but I’ve already finished the note, and I’ve already registered on that machine. 然后那人把我的纸收了就直接贴入境单了，很丝滑很快。这次不去京都，就直接走南海电车了，更便宜。坐到 namba 才 900。</p><p>吃完饭，就去附近的 apple 店买个表。现在日本的 apple 店居然不能退税了，只有部分非官方 retailer 可以退税，但是我又不知道哪里靠谱。46mm 的要 3200，国内才 2600。但是因为国内啥啥都没，没快充慢的一批，所以还是买了。总不能为这个专门跑一次 hk 吧。</p><p>去 711 给交通卡充值，发现只能用现金。幸亏我带了 15000 jpy 来，本意是上次玩没用掉，这次结果救了命。</p><p>晚上出去吃饭回来，发现大阪真的冷，幸亏穿了两件，不然真的冻死。</p><p>回来发现日本的窗子真的好隔音，薄薄的平开窗，毫无特点的铝合金，居然这么猛。</p><h1 id="D2-京都"><a href="#D2-京都" class="headerlink" title="D2 京都"></a>D2 京都</h1><p>因为地铁卡充值要现金感觉不方便，所以我今天就打算不用地铁卡，用 iPhone Wallet 了。这鬼东西又不能用信用卡，我弄了半天才弄了个储蓄卡上去。上地铁又刷坏了，墨迹了半天。到了大阪梅田站，她出不了站，刷了半天不知道怎么就出去了。结果到了 JR 大阪站发现进不去，问工作人员说这个是 osaka subway，我觉得很奇怪，因为 icoca 是通用的，工作人员也说我的 iPhone 是可以的。于是我觉得肯定是被锁了。无论如何，这里可以支付宝买纸票，于是我们就去京都了。</p><p>京都的公交车排队排了感觉有半个多小时，后面上了个临时的加班车。这边还不让带行李上公交车，幸亏我们都放在了大阪，就带了个小包。公交车里面非常拥挤，但我们进去的早，所以有个位置。公交车开的特别慢，感觉甚至不如走路。开到清水寺附近感觉花了大概二十多分钟。我们试图在五年阪下车，结果被堵住了，然后司机就不让我们下了。然后后面一个傻逼老外就说 you come here late, you have to follow their rules, it’s not your country 啥的。非常 offensive。</p><p>清水寺我们没进去，我对象说没啥意思，我之前去过，感觉也没啥意思，人还特别多。门口的御守他也觉得贼丑。然后我们就顺着三年半二年阪往下走，去找法观寺。我们都已经看到高台寺公园了，发现法观寺走过了，有绕回去找。结果法观寺就是我们来的路上的那个我觉得一般的唐代风格的塔，好像是京都的最老的塔，也不让进去。然后我们又走回到高台寺。</p><p>高台寺的垂樱是挺好看的。</p><p>高台寺出来很快就到了八阪神社，里面和好吃街一样，我们找了下垂樱在哪里，就去吃那个鳗鱼饭了。</p><p>鳗鱼饭吃完出来，就在木屋町那条小路那边拍樱花，感觉比鸭川的樱花好看。</p><p>然后又绕回去看了下花见小路，照旧非常无聊。路尽头是什么建仁寺的，大家都没有兴趣看。</p><p>然后又走了一大段路去看顶法寺。顶法寺不要钱，里面的樱花很漂亮。还有几个小和尚的雕塑也很有意思。</p><p>晚上实在走不动了，就坐了京阪电车，因为阪急还要多走 300m。至于京都 JR，因为坐公交体验太差，完全不考虑了。我们坐的是 18.59 的京阪电车去的大阪。其实以后真的可以坐京阪电车到京都东边，比 JR 方便很多，还便宜。但这次京阪电车属实是个坑货。它终点站在大阪是 yodoyabashi 淀屋桥，但是又有一些车是到一个叫 yodo 淀的地方。然后我就被这个车丢在了 yodo。后面上了个准急的，几乎就是站站乐了，感觉总共花了大概一个多小时才到大阪。</p><p>回到大阪，找大阪地铁说明了情况，列车长问我们要不要 refund。我说不要，他就说 OK。然后就解锁了，结果发现之前坐的 namba 的南海的钱居然也没被扣。南海真的血亏啊。</p><h1 id="D3-姬路"><a href="#D3-姬路" class="headerlink" title="D3 姬路"></a>D3 姬路</h1><p>起来去姬路。这次发现他们电车同一个方向有两个轨道，一边是 local 一边是 express。上车前可以看站台上的屏幕，上车时可以看车的屏幕确认。上车前可以看 Google 或者 Apple 的 map 到达的时间确定自己是不是对应班次的车。上车后也可以听播报。基本上 Rapid、Express 的车都比较推荐，大阪、神户市区的大站基本都停。Limited Express 特急要特价券我没见过。</p><p>姬路是个小城市，我们的酒店在车站南边一点。应该是此行中比较大的酒店了。酒店不能提前办理入住，但是可以预先寄存。</p><p>姬路站北面正对姬路城，通过一个大道可以直接走过去。大道两侧不少店铺比较出名，我们吃了个咖啡，然后就立即前往姬路城了。</p><p>城里面樱花很漂亮，右边走还有个动物园。登城口在左边，要排队，但是队伍很快，等了大概半小时就进去了。买票 +50 yen 就能得到一个旁边的花园的门票。花园里面有个餐厅，我们去的时候已经关门了。姬路城主要就是两块，一个是西之丸庭院，可以脱掉鞋子登上去，类似于一个走廊，有多个城橹。从化妆橹可以下来，据说这个是给城主夫人化妆用的地方。西之丸庭院相比大阪城比较小巧，里面也有不少樱花。</p><p>从庭院出来就可以走到天守的口了。然后就是穿过一道道什么 yi 之门、wa 之门、ni 之门，然后走到天守的下面的小庭院内。然后就走过水之 X 门走到大天守里面。大天守有 6 层，第 2 层开始大排长队，人贼多。</p><h1 id="D4-神户"><a href="#D4-神户" class="headerlink" title="D4 神户"></a>D4 神户</h1><p>神户的酒店在三宫门口，应该是本次我找的最近、最便宜的酒店了，才 500。住宿体验很舒适。</p><p>放完东西，去神户动物园。我们应该是顶门到的，进去刷票的时候，我对象把票根掉了，检票员说必须要把票找回来。</p><p>在那个破落商店里面有一家后来知道叫 Yellow Submarine 的桌游超市，还挺大的。回来我去问了下有没有那个骰子游戏，他带我去找，然后翻了下，说 sold out 了。我以为附近南京町还有一家，走过去发现那是个别的地方，已经关门了。</p><h1 id="D5-奈良-大阪"><a href="#D5-奈良-大阪" class="headerlink" title="D5 奈良 - 大阪"></a>D5 奈良 - 大阪</h1><p>下了奈良站，在 5 号口附近找到一个柜子，600 yen 就可以放两个小箱子和一个包了。不过只接受 100 yen 的硬币，我还要去旁边换。</p><p>出去吃了那个口水麻薯，一堆老外在那边拍照。去看了兴福寺，要钱，没意思没进去。兴福寺下来就能看到鹿，鹿很现实，看到我们没买饼就不磕头了。我们到最后也没买，主打白嫖。</p><p>旁边的奈良博物馆关门了。</p><p>我印象里上次去若草山有个大坡可以坐着休息，但这次去好像就是一些大草坪可以走，有椅子可以坐。大草坪也许是封起来了吧，因为当时看到大草坪上没有人，而且往若草山方向有被拦住。</p><p>在大阪这最后一晚住的酒店是最拉胯的。</p><h1 id="D6-大阪"><a href="#D6-大阪" class="headerlink" title="D6 大阪"></a>D6 大阪</h1>]]></content>
    
    
    <summary type="html">&lt;p&gt;趁着清明节又去了一趟关西。本以为是度假，但实际上累得要死。&lt;/p&gt;</summary>
    
    
    
    
    <category term="游记" scheme="http://www.calvinneo.com/tags/游记/"/>
    
  </entry>
  
</feed>
