文章

CommonJS 规范与实现深度解析

CommonJS 三要点:同步加载、模块缓存(同一文件只执行一次)、值拷贝导出(导出值不随模块内变化更新)。 理解这三点就能解释 Node 模块的几乎所有行为。

CommonJS 规范与实现深度解析

一句话概括

CommonJS 三个字记一辈子:同步加载(require 读完文件执行完才往下走)、模块缓存(同一文件只跑一次,后续返回缓存)、值拷贝导出(导出的原始值不会随模块内部变化而更新)。这三点解释了 CJS 的全部行为。

核心知识点

1. 每个文件是独立模块——靠 IIFE 实现

1
2
3
4
5
6
7
8
9
// math.js(导出)
const add = (a, b) => a + b;
const SECRET = '外面拿不到';       // 没挂 exports,天然封装
module.exports = { add };

// app.js(导入)
const math = require('./math');
math.add(1, 2);    // 3
math.SECRET;       // undefined —— 模块作用域自动隔离

Node.js 把每个文件包在一个函数里执行:

1
2
3
(function(exports, require, module, __filename, __dirname) {
  // 你的代码在这里
});

顶层变量全是这个函数的局部变量,默认不外泄——这比其他语言靠 private 关键词实现封装更优雅。

2. module.exports vs exports — 经典坑

1
2
3
4
5
6
7
8
// Node 内部隐式执行了这一行:
// const exports = module.exports;

exports.foo = 'bar';        // ✅ 改的是同一对象的属性
exports = { foo: 'bar' };   // ❌ exports 指向新对象,module.exports 没变!

// 🚨 铁律:永远只用一种写法
module.exports = { add };   // ✅ 永不出错

一句话出师: require() 永远只认 module.exports,exports 只是它的快捷引用。一旦给 exports 赋新值,引用断了,此后加什么都不出现在 require 结果里。

3. 缓存机制——代码只跑一次

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
// counter.js
let count = 0;
module.exports = {
  count,                   // 导出的是当前值 0 的快照
  inc: () => { count++; },
  getCount: () => count    // 闭包访问内部变量,能拿到最新值
};

// app.js
const c1 = require('./counter');
const c2 = require('./counter');
console.log(c1 === c2);      // true —— 单例

c1.inc();
console.log(c2.getCount());  // 1  —— 闭包共享了内部 count
console.log(c1.count);       // 0  —— count 字段是值快照,不更新!

4. require 查找链路

1
2
3
4
// 优先级从高到低:
require('fs');           // ① 核心模块(Node 内置,最快)
require('./utils');      // ② 相对/绝对路径 → 补全 .js → .json → .node → 当目录找 /index.js
require('lodash');       // ③ node_modules → 从当前目录逐级向上冒泡,直到 /

5. require 简易实现——看懂本质

1
2
3
4
5
6
7
8
9
10
11
12
13
14
function myRequire(filePath) {
  const filename = path.resolve(filePath);              // 1. 解析绝对路径
  if (cache[filename]) return cache[filename].exports;  // 2. 走缓存

  const module = { exports: {} };
  cache[filename] = module;                             // 3. ⚠️ 先缓存(为循环依赖留后门)

  const wrapper = `(function(exports, require, module, __filename, __dirname) {
    ${fs.readFileSync(filename, 'utf8')}
  })`;
  eval(wrapper).call(module.exports, module.exports, myRequire, module, filename, dirname);

  return module.exports;                                // 4. 返回
}

第 3 步是理解循环依赖的钥匙: 缓存设在执行代码之前——先占位 {},代码跑完再填充。循环依赖方拿到的是半成品空对象而不是 undefined。

其实你每天都在用

  • require('fs')、require('path'):Node 标准库全是 CJS
  • Webpack __webpack_require__:在浏览器里山寨了一套 CJS 运行时
  • Babel / tsc 转译:你写的是 ESM,Node 里实际跑的是被转成的 CJS
  • .eslintrc.js / jest.config.js / webpack.config.js:Node 端配置文件默认 CJS
  • npm 包的 "main" 字段:指向 CJS 入口,require('pkg') 就找它

常见误解(FAQ)

  • ❌ 误区:「exports 和 module.exports 一样,随便用」 起点一样(exports === module.exports),但一旦任一方被重新赋值就断了。铁律:只用 module.exports,或只往 exports 上加属性,别混用赋值操作。

  • ❌ 误区:「CJS 导出后模块内部改了值会同步到外面」 不会。module.exports = { count } 拷贝的是赋值那一刻的值。想要”实时值”得用 getter:Object.defineProperty(module.exports, 'count', { get: () => count })。这和 ESM 的 Live Binding 是最本质的区别。

  • ❌ 误区:「循环依赖必须完全避免」 大项目完全避免不现实。CJS 能处理——先缓存再执行,但拿到的模块是半成品({})。解法:把对被依赖方的取值延迟到函数调用时(运行时 require),别在模块顶层立即取值。

  • ❌ 误区:「require 和 import 只是语法糖差异」 require 运行时同步读文件 + 立即执行;import 编译时静态解析结构,执行时再取值。这决定了 CJS 能条件 require(if(x) require('a'))、ESM 能 Tree Shaking——设计哲学完全不同。

一句话总结

同步读文件、缓存 module.exports、值拷贝导出——CommonJS 三句话讲完,正因为简单,npm 几十万个包才敢拿它当基石。

本文由作者按照 CC BY 4.0 进行授权