Bug Tracker
This document records known bugs of the ErisPulse SDK and their repair status, arranged in chronological order by the time of repair.
For Readers No software is perfect from the start, and even the most careful developers can leave small mistakes. This tracker only includes issues that have practical impact on operation—those that are too minor to even reach the "minor" level will not appear here. Although the list contains many entries marked as "severe," the original intention of publicly documenting these bugs is to make troubleshooting and tracing smoother, not to create anxiety: problems that are visible, recorded, and repaired are themselves proof that the project is constantly improving. Don't worry when you see this list; it is a troubleshooting tool, not a source of fear.
How to Read & Maintenance Conventions
- Each bug record contains structured fields such as problem description, root cause analysis, affected version range, repair solution, etc. It is recommended to check the "affected version" before upgrading to see if it covers the version currently in use.
- If you need to add a new bug entry, please supplement content at the corresponding location, following the field specifications and severity/type classification described below.
Field Description
Mandatory Fields
| Field | Description |
|---|---|
| Problem | The external manifestation of the bug, the abnormal phenomenon observable by the user. Try to give error messages or typical scenarios |
| Root Cause | Root cause analysis, pointing to specific code defects (including "root cause chain" diagrams for complex scenarios) |
| Affected Version | Affected version range, format introduced version - fixed version (including both dev versions) |
| Fixed Version | The specific version number that fixed the bug |
| Repair Content | Brief description of the repair solution, including key code changes |
| Repair Date | The release date of the corresponding fixed version, using the YYYY/MM/DD format |
| Severity | Marked according to the "Severity Classification" below |
| Type | Marked according to the "Type Classification" below, can be combined (e.g., Adapter / Router) |
Optional Fields
| Field | Description | Applicable Scenarios |
|---|---|---|
| Reproduction Steps | The minimal reproducible path to trigger the bug | Complex bugs, sporadic bugs are recommended to supplement |
| 关联 | Related Issue / PR / Commit links | Supplement when there is external discussion record |
| Regression Test | Test case location for verifying repair and preventing regression | Supplement when corresponding pytest cases have been written |
Severity Classification
| Identifier | Level | Judgment Standard | Typical Manifestation |
|---|---|---|---|
| 🔴 | Severe | Causes process crash, data loss/damage, complete unavailability of core functions, security vulnerabilities | OOM Kill, message cannot be sent, module cannot be loaded, hot reload failed |
| 🟡 | Medium | Function abnormality but with workaround, non-core function failure, sporadic problems | Status judgment error, repeated trigger, cache expiration, inaccurate error message |
| 🟢 | Minor | Does not affect core functions, only code quality or experience issues, potential risks have not occurred | Deprecated API, dead code, missing warning logs |
Type Classification
| Type | Scope |
|---|---|
| Configuration System | ConfigManager, configuration read/write, configuration Schema, hot update |
| Event System | Event module (command/message/notice/request/meta), event distribution, handler registration |
| Adapter | AdapterManager, BaseAdapter, account parsing, Bot status, middleware |
| Router | RouterManager, HTTP/WebSocket/SSE routing, rate limiting, CORS |
| Client | HttpClient, ClientWebSocket, aiohttp wrapper |
| Storage | StorageManager, SQLite, SQL builder, nested keys |
| Loading System | Loader, LazyModule, ModuleInitializer, strict mode, module discovery |
| CLI | epsdk command, init/run/install, parameter parsing, signal handling |
| Runtime | sdk.run/restart/uninit, lifecycle, signal, subprocess |
Entry Template
When adding a new bug entry, follow the following format:
### [BUG-XXX] Title
**Problem**: Problem description (error message or typical phenomenon)
**Root Cause**: Root cause analysis
**Affected Version**: Introduced version - Fixed version
**Fixed Version**: x.x.x
**Repair Content**: Repair solution
**Repair Date**: YYYY/MM/DD
<!-- Optional fields -->
**Reproduction Steps**: (Recommended for complex bugs)
**关联**: (Issue/PR links)
**Regression Test**: (Test case path)
**Severity**: 🔴 Severe | 🟡 Medium | 🟢 Minor
**Type**: Configuration System / Event System / Adapter / Router / Client / Storage / Loading System / CLI / Runtime
Statistical Overview
| Severity | Quantity |
|---|---|
| 🔴 Severe | 16 |
| 🟡 Medium | 18 |
| 🟢 Minor | 3 |
| Total | 37 |
| Type | Quantity |
|---|---|
| Adapter | 6 |
| Configuration System | 11 |
| Event System | 7 |
| CLI | 3 |
| Storage | 3 |
| Loading System | 3 |
| Router | 2 |
| Client | 1 |
| Runtime | 1 |
Note: A single bug can belong to multiple types, the above table counts by main type.
Note: BUG-028 / BUG-031 numbering is missing (abandoned during registration, numbering is not recycled to maintain stability).
Fixed Bugs
[BUG-001] Event handler registration repetition leads to event being processed multiple times
Problem: When using multiple @message / @notice decorators to register handlers, the same event is triggered multiple times, causing commands to be executed multiple times and logs to be output repeatedly.
Root Cause: BaseEventHandler registers handlers to the adapter event bus without deduplication logic, each decorator mounts once to the bus, and events are called multiple times during distribution.
Affected Version: 2.2.0-dev.0 - 2.2.1-dev.0
Fixed Version: 2.2.1-dev.0
Repair Content: Optimize BaseEventHandler to ensure each event type registers only once to the adapter, avoiding repeated triggers.
Repair Date: 2025/08/18
Severity: 🔴 Severe
Type: Event System
[BUG-002] Adapter configuration path type error in Init command
Problem: When using the ep init command for interactive initialization, selecting the configuration adapter results in a type error:
Interactive initialization failed: unsupported operand type(s) for /: 'str' and 'str'
Root Cause: When adjusting the configuration file path in version 2.3.7, the parameter type is inconsistent. Path operators are used internally in _configure_adapters_interactive_sync, but the method receives str type parameters.
Affected Version: 2.3.7 - 2.3.9-dev.1
Fixed Version: 2.3.9-dev.1
Repair Content: Change the parameter type of _configure_adapters_interactive_sync from str to Path, and pass Path objects directly during calls.
Repair Date: 2026/03/23
Severity: 🟡 Medium
Type: CLI
[BUG-003] Commands fail after restart
Problem: After calling sdk.restart(), commands registered via @command cannot be triggered, resulting in the robot not responding when sending commands.
Root Cause: adapter.shutdown() clears the event bus, but the _linked_to_adapter_bus status is not reset to False, causing the _process_event method to think it has been mounted to the adapter bus and skip re-mounting.
Affected Version: 2.2.x - 2.4.0-dev.2
Fixed Version: 2.4.0-dev.3
Repair Content: Introduce _linked_to_adapter_bus status tracking, and after clear_handlers() disconnects the bus, re-mount automatically on the next register() to adapt to shutdown/restart scenarios.
Repair Date: 2026/04/09
Severity: 🔴 Severe
Type: Event System
[BUG-004] Lifecycle event handlers not cleaned up
Problem: After sdk.restart(), old lifecycle event handlers still exist and are triggered repeatedly, causing the same event to be processed multiple times.
Root Cause: The lifecycle._handlers dictionary is never cleared in uninit(), so old handlers and new handlers coexist after restart.
Affected Version: 2.3.0 - 2.4.0-dev.2
Fixed Version: 2.4.0-dev.3
Repair Content: Clear lifecycle._handlers at the end of the Uninitializer cleanup process (after all events are submitted), to ensure old handlers are removed.
Repair Date: 2026/04/09
Severity: 🟡 Medium
Type: Runtime
[BUG-005] Event.is_friend_add/is_friend_delete detail_type inconsistent with OB12 standard
Problem: Event.is_friend_add() checks detail_type == "friend_add", Event.is_friend_delete() checks detail_type == "friend_delete", but the OneBot12 standard defines detail_type values as "friend_increase" and "friend_decrease". This is inconsistent with the values used in notice.py's on_friend_add/on_friend_remove decorators, causing handlers registered via decorators to trigger, but the corresponding is_friend_add()/is_friend_delete() judgment methods return False.
Root Cause: wrapper.py uses non-standard naming, while notice.py uses the correct OB12 standard naming.
Affected Version: Implemented to date
Fixed Version: 2.4.2-dev.1
Repair Content: Change is_friend_add()'s matching value from "friend_add" to "friend_increase", and is_friend_delete() from "friend_delete" to "friend_decrease".
Repair Date: 2026/04/13
Severity: 🟡 Medium
Type: Event System
[BUG-006] adapter.clear() does not clean up _started_instances, causing incorrect status after restart
Problem: The AdapterManager.clear() method clears _adapters, _adapter_info, handlers, and _bots, but omits _started_instances. If the adapter is running and clear() is called, _started_instances retains dangling references, causing incorrect status judgment after restart.
Root Cause: When _started_instances was introduced in 2.4.0-dev.1, it was not cleared synchronously in clear().
Affected Version: 2.4.0-dev.1 - 2.4.2-dev.0
Fixed Version: 2.4.2-dev.1
Repair Content: Add self._started_instances.clear() in the clear() method.
Repair Date: 2026/04/13
Severity: 🟡 Medium
Type: Adapter
[BUG-007] command.wait_reply() uses deprecated asyncio.get_event_loop()
Problem: The CommandHandler.wait_reply() method uses asyncio.get_event_loop() to create futures and obtain timestamps, which has been deprecated in Python 3.10+. In asynchronous contexts, asyncio.get_running_loop() should be used. This is inconsistent with the get_running_loop() used in the wrapper.py wait_for() method.
Root Cause: The code was developed using the old API, and the newly added wait_for() method used the correct API but did not retrospectively fix the old code.
Affected Version: 2.3.0-dev.0
Fixed Version: 2.4.2-dev.1
Repair Content: Replace asyncio.get_event_loop() in command.py with asyncio.get_running_loop() in two places.
Repair Date: 2026/04/13
Severity: 🟢 Minor
Type: Event System
[BUG-008] Bot offline events are repeatedly submitted during shutdown
Problem: When calling adapter.shutdown() to close all adapters, _update_bot_status() repeatedly submits Bot offline events during the shutdown process, causing the same batch of Bots to be marked offline multiple times and triggering the adapter.bot.offline lifecycle event multiple times.
Root Cause: The Bot status tracking system introduced in 2.4.0-dev.1 did not set a "shutting down" flag during shutdown(), so _update_bot_status() could not distinguish between normal offline and cascading offline during shutdown.
Affected Version: 2.4.0-dev.1 - 2.4.2-dev.1
Fixed Version: 2.4.2-dev.1
Repair Content: Add _is_being_shutdown flag in AdapterManager, set to True at the start of shutdown() and cleared at the end; _update_bot_status() skips repeated submissions during shutdown after checking the flag.
Repair Date: 2026/04/21
Severity: 🟡 Medium
Type: Adapter
[BUG-009] Synchronous access to BaseModule in LazyModule causes incomplete initialization
Problem: When users access the attributes of a lazy-loaded BaseModule in a synchronous context, the module uses loop.create_task() for asynchronous initialization but does not wait, leading to race conditions when attributes are accessed before initialization is complete.
Root Cause: _ensure_initialized() uses loop.create_task(self._initialize()) and returns immediately, without ensuring initialization is complete.
Affected Version: 2.4.0-dev.0 - 2.4.2-dev.1
Fixed Version: 2.4.2-dev.2
Repair Content: In synchronous contexts, BaseModule initialization is changed to use asyncio.run(self._initialize()) to ensure initialization is complete before returning. The transparent proxy feature is maintained, and users do not need to be aware of synchronous/asynchronous differences.
Repair Date: 2026/04/21
Severity: 🟡 Medium
Type: Loading System
[BUG-010] Multi-threaded configuration writes cause data loss
Problem: In a multi-threaded environment, when multiple threads simultaneously call config.setConfig(), the _flush_config() read-modify-write operation is not atomic, potentially causing some writes to be lost.
Root Cause: Although _flush_config() uses RLock, there is no file lock protection between file reads and writes, and the _schedule_write Timer may be triggered multiple times, causing overwrites.
Affected Version: 2.3.0 - 2.4.2-dev.1
Fixed Version: 2.4.2-dev.2
Repair Content:
- Add file lock mechanism (
_file_lock) to ensure atomic file operations - Use temporary files for writing and atomically rename (
os.replace/os.rename) - Improve
_schedule_writeTimer cancellation and rescheduling logic
Repair Date: 2026/04/21
Severity: 🔴 Severe
Type: Configuration System
[BUG-011] Windows Ctrl+C cannot stop the program
Problem: When running python main.py directly on Windows, pressing Ctrl+C does not terminate the program. The program starts normally and outputs the routing server information, but Ctrl+C has no response and can only be forcibly terminated through Task Manager. However, it can be stopped normally when started via epsdk run—but epsdk run runs through a subprocess model.
Root Cause: The Hypercorn ASGI server's serve() function internally registers its own SIGINT handler via signal.signal(SIGINT, handler), overriding Python's default KeyboardInterrupt handling mechanism. When Hypercorn is started as a background task via asyncio.create_task(), its internal shutdown process cannot be triggered properly (because it expects the worker_serve mode), causing the Ctrl+C signal to be swallowed by Hypercorn without triggering any cleanup actions.
Affected Version: 2.3.6 - 2.4.2
Fixed Version: 2.4.3-dev.0
Repair Content:
- Switch the ASGI server from Hypercorn to Uvicorn (
pyproject.tomldependency change) - Use
uvicorn.Server._serve()to start the server directly, bypassing thecapture_signalssignal handling context manager - Implement graceful shutdown via
server.should_exit = True, canceling the background task if timeout occurs - Synchronously remove the subprocess running model and
runtime/cleanup.pycleanup module (no longer needed for subprocess cleanup)
Repair Date: 2026/04/28
Severity: 🔴 Severe
Type: CLI / Runtime
[BUG-012] Hot restart after updating modules does not take effect
Problem: After executing sdk.restart() soft restart, the new code (such as new API routes) of modules/ adapters upgraded via epsdk install does not take effect, and the old logic is still running. The latest code can only be loaded by completely restarting the process.
Root Cause: _do_restart() calls entry_point.load() during reinitialization, but this function returns a cached module object from sys.modules instead of reloading from disk.
Affected Version: Early versions - 2.4.3-dev.1
Fixed Version: 2.4.3-dev.1
Repair Content: Clean the cache of loaded modules/adapters in sys.modules before init() and after uninit(), so that entry_point.load() reloads the latest code from disk. Added _collect_top_level_modules() and _invalidate_module_cache() helper methods, deriving top-level module names through top_level.txt or entry-point value.
Repair Date: 2026/05/03
Severity: 🔴 Severe
Type: Loading System / Runtime
[BUG-013] Module loading strategy sorting logic error
Problem: ModuleLoadStrategy provides a priority field to declare the initialization priority of modules, but the implementation of the loading strategy has an error, causing modules not to be initialized in the expected priority order, but rather in the default order of entry_points(). When modules have initialization dependencies, the correct initialization order cannot be ensured through priority.
Root Cause: The implementation of the sorting logic in the loading strategy is incorrect, and initialize_modules() does not sort the module list by priority.
Affected Version: 2.3.4 - 2.4.5-dev.2
Fixed Version: 2.4.5-dev.3
Repair Content: Before the initialize_modules() traversal, sort the module list by priority in descending order. Modules with the same priority maintain their original relative order (stable sorting).
Repair Date: 2026/05/15
Severity: 🟡 Medium
Type: Loading System
[BUG-014] Adapter middleware returning None causes event data loss
Problem: When adapter.emit() executes the OneBot12 middleware chain, if a middleware returns None (for example, forgetting to return data), the subsequent middleware and all event handlers receive processed_data as None, causing event processing to fail completely.
Root Cause: The middleware chain implementation processed_data = await middleware(processed_data) does not check if the return value is None, directly overwriting the result of the previous step.
Affected Version: unknown - 2.4.5-dev.3
Fixed Version: 2.4.5-dev.4
Repair Content: If the middleware returns None, ignore the return value, retain the original data, and output a warning-level log.
Repair Date: 2026/05/15
Severity: 🔴 Severe
Type: Adapter / Event System
[BUG-015] Configuration file path depends on working directory
Problem: The configuration file path in ConfigManager is a relative path "config/config.toml" by default, which depends on os.getcwd() for resolution at runtime. If the working directory changes during runtime (for example, by using os.chdir()), configuration file read/write operations will point to the wrong location, causing configuration loss or reading old data.
Root Cause: __init__ directly stores relative paths without resolving them to absolute paths at initialization.
Affected Version: 2.3.7 - 2.4.5-dev.3
Fixed Version: 2.4.5-dev.4
Repair Content: In ConfigManager.__init__(), if the passed path is a relative path, automatically resolve it to an absolute path using os.path.abspath().
Repair Date: 2026/05/15
Severity: 🟡 Medium
Type: Configuration System
[BUG-016] BaseStorage confuses storing value None with key not existing
Problem: BaseStorage.get_multi() / __getattr__() cannot distinguish between "key not existing" and "key's value is None", and users explicitly storing None will be treated as "key not existing" when read again.
Root Cause: The value retrieval logic directly uses value is None to determine if the key exists, lacking an independent "missing" marker.
Affected Version: Early versions - 2.4.6-dev.6
Fixed Version: 2.4.6-dev.6
Repair Content: Introduce _SENTINEL sentinel value to distinguish "key not existing" from "value is None", so they are no longer confused.
Repair Date: 2026/06/07
Severity: 🟡 Medium
Type: Storage
[BUG-017] WebSocket route auto_accept flag lost after service restart
Problem: After a service restart (such as sdk.restart()), the auto_accept configuration of all WebSocket routes becomes False, and the originally expected auto-accept connections become suspended, causing the client to not receive responses for a long time, manifesting as a WS connection hang.
Root Cause: _restore_routes_from_records() hardcodes auto_accept as False when restoring routes from persistent records, without reading the value from the original record; also, when the route storage tuple was extended from a binary tuple to a ternary tuple, the restoration logic was not synchronized.
Affected Version: 2.3.8-dev.0 - 2.4.6-dev.6
Fixed Version: 2.4.6-dev.6
Repair Content: The route storage tuple is extended to (handler, auth_handler, auto_accept), and _restore_routes_from_records() reads the real auto_accept value from the record instead of hardcoding False.
Repair Date: 2026/06/07
Severity: 🔴 Severe
Type: Router
[BUG-018] Concurrent calls to HTTP/WS client cause crashes and connection leaks
Problem: The HTTP and WebSocket clients in Core/client.py have multiple stability defects in concurrent scenarios, leading to connection leaks or process crashes:
- Concurrent calls to
ClientWebSocket.receive()by multiple coroutines cause aiohttp to throwConcurrent call to receive() is not allowed _get_http_session()/_get_ws_session()concurrent calls may create multiple sessions, and_drain_sessions()does not close old connections, causing connection leaks- The exception handling order in
request()is incorrect:except ClientConnectionError(ErisPulse exception) is never triggered, aiohttp connection errors are caught by the generalexcept Exception, causing the "retry + session rebuild" logic (dead code) to never execute send_json()ignores themode="binary"parameter;_get_ws_session()does not pass default request headers
Root Cause: The initial client implementation (2.4.6-dev.5) lacked concurrent protection and improper handling of aiohttp exception hierarchy and ErisPulse custom exception inheritance.
Affected Version: 2.4.6-dev.5 - 2.4.8
Fixed Version: 2.4.8
Repair Content:
- Add
_recv_lockto serialize allreceive()/receive_text()/receive_bytes()calls - Add
_session_lockto protect session creation;_drain_sessions()is changed to an asynchronous method and truly closes old sessions - Refactor
request()exception handling order:asyncio.TimeoutError→aiohttp.ClientConnectionError(triggers session rebuild) →aiohttp.ClientError→ClientError(transparently passed) →Exception - Fix
send_json()mode handling,_get_ws_session()default request header passing,close()concurrent race condition,HttpResponse.__aexit__repeatedrelease()
Repair Date: 2026/06/12
Severity: 🔴 Severe
Type: Client
[BUG-019] Adapter hot reload causes route conflicts and reload failure
Problem: When third-party modules (such as Dashboard) trigger adapter hot reload, or when adapter startup fails and retries, old routes (such as onebot11_default) are not cleared, causing WebSocket path ... already registered conflicts and reload failures. A complete process restart is required to restore.
Root Cause: AdapterManager.shutdown() only clears routes by unregister_all_by_namespace(platform), but adapters (such as OneBot11) register WS routes with onebot11_{account_name} as the namespace, resulting in a granularity mismatch and making the cleanup an empty operation; startup failure retry routes are also not cleared of previous residual routes.
Affected Version: Early versions - 2.4.9
Fixed Version: 2.4.9
Repair Content:
- Route registration automatically tracks
owner → namespaceownership relationships throughcurrent_ownerContextVar - Add
unregister_all_by_owner(owner), cleaning up by owner during stop/restart, covering fine-grained namespaces - Add
_stop_adapter(platform)primitive (stop equals cleanup), binding stopping adapter and reclaiming its registered resources in a single call,restart()and startup failure retry both go through this entry - Add framework-level
adapter.restart(platform)API, third-party modules should call this method instead of directly operating adapter instances
Repair Date: 2026/06/12
Severity: 🔴 Severe
Type: Adapter / Router
[BUG-020] Subprocess mode ep run <script> cannot find subpackages in script's directory
Problem: When running a script with ep r .\main.py in non-hot-reload mode, if the script has relative imports (such as from qg import ...), it reports No module named 'qg' error. While --reload mode works normally.
Root Cause: Non-hot-reload mode directly calls runpy.run_path() to execute the script, which does not automatically add the script's directory to sys.path. While --reload mode runs via subprocess.Popen subprocess, the subprocess automatically inherits the current working directory, and sys.path[0] is the script's directory, so it works normally.
Affected Version: 2.5.0 - 2.5.2-dev.0
Fixed Version: 2.5.2-dev.0
Repair Content: Before calling runpy.run_path(), manually insert the script's directory into sys.path[0].
Repair Date: 2026/06/27
Severity: 🟡 Medium
Type: CLI
[BUG-021] SQL query builder rejects valid wildcard and list expressions
Problem: The _build_select_sql() of SQLiteQueryBuilder validates all SELECT columns with _validate_identifier(), which uses a strict whitelist regex ^[a-zA-Z_][a-zA-Z0-9_]*$, causing legitimate SQL syntax to be incorrectly judged as unsafe column names:
SELECT *—*is a standard SQL wildcardSELECT COUNT(*)— aggregate functionSELECT users.name— qualified column nameSELECT col AS alias— column alias
Among them, Select("*") is used by modules like Cron, causing module on_load execution failure and module loading failure.
Root Cause: In version 2.4.6, SQL injection protection was enhanced, introducing _validate_identifier() whitelist validation. This validation is applied to all column names, but does not distinguish between read端 (SELECT/ORDER BY) and write端 (INSERT/UPDATE). SELECT columns allow complex SQL expressions and should not be restricted by simple identifier whitelist.
Affected Version: 2.4.6 - 2.5.2-dev.1
Fixed Version: 2.5.2-dev.2
Repair Content: Change the SELECT/ORDER BY column validation from whitelist mode to blacklist mode:
- Add
_validate_select_column()function, only blocking SQL injection dangerous characters (;'"--/**/\x00newline) - Allow any valid SQL column expression (
*,table.*,table.column,COUNT(*),col AS alias, etc.) - INSERT/UPDATE column names still maintain strict whitelist validation (only allow simple identifiers)
Repair Date: 2026/06/29
Severity: 🔴 Severe
Type: Storage
[BUG-022] _resolve_account() account resolution regression (_accounts_data not populated)
Problem: After the 2.5.2 configuration system refactoring, multi-account adapters declared with AccountConfigClass report ValueError("未声明 AccountConfigClass,无法解析账户") when calling methods like wait_reply, reply that require sending messages. Even if the adapter correctly configures multi-account information, account resolution still fails.
Root Cause: In 2.5.2-dev.5, _load_accounts() (responsible for reading configuration + validation + populating _accounts_data) was refactored into _ensure_accounts_exist() (only generates configuration template), but _resolve_account() still checks self._accounts_data is None. Since _ensure_accounts_exist() no longer populates _accounts_data, this attribute remains None, causing _resolve_account() to prematurely return (None, None), account resolution fails completely.
Root Cause Chain:
_load_accounts() was deleted
→ __init__ no longer populates _accounts_data
→ _accounts_data is always None
→ _resolve_account() checks _accounts_data is None → return (None, None)
→ downstream calls to _resolve_account (e.g. call_api) get None
→ triggers error
Affected Version: 2.5.2-dev.5 - 2.5.2
Fixed Version: 2.5.3
Repair Content: In BaseAdapter.__init__, after _ensure_accounts_exist(), restore the population of _accounts_data:
if self.AccountConfigClass is not None:
self._ensure_accounts_exist()
self._accounts_data = self.accounts # restore population, data source is real-time read accounts attribute
The _resolve_account() logic remains unchanged, fully backward compatible:
- Adapters that do not declare
AccountConfigClass:_accounts_dataremainsNone→ return(None, None) - Adapters that declare
AccountConfigClass:_accounts_datais populated → normal resolution - Adapters that override
_load_accountsor manually set_accounts_data: override aftersuper().__init__()call, highest priority
Repair Date: 2026/07/07
Severity: 🔴 Severe
Type: Adapter / Configuration System
[BUG-023] After modifying account configuration, adapter cache is not refreshed, causing account resolution failure
Problem: After users modify the multi-account adapter's account configuration (such as filling in the token) through the Dashboard, the adapter still uses the old cache, and calling message-sending-related methods reports 未找到可用账户 (account_id=default). The process must be restarted for the new configuration to take effect.
Root Cause: _accounts_data is only read from the configuration storage once at BaseAdapter.__init__, and is never refreshed afterwards. AdapterManager._run_adapter() and restart() do not re-read the account configuration before calling adapter.start(), causing the cache to be out of sync with the actual configuration.
Affected Version: 2.4.6 - 2.5.4
Fixed Version: 2.5.4
Repair Content: In AdapterManager._run_adapter() and restart(), refresh adapter._accounts_data = adapter.accounts before calling adapter.start(), ensuring that the latest configuration is used each time the adapter starts.
Repair Date: 2026/07/09
Severity: 🔴 Severe
Type: Adapter / Configuration System
[BUG-024] storage.set() writing large number ID keys triggers OOM Kill
Problem: When calling storage.set() to write a nested key path containing a large numeric field (such as QQ group number 871684833), the process is killed by container OOM (exit code -9), and the service crashes directly without recovery.
Root Cause: In the recursive implementation of _set_nested_value, the pure numeric field in the nested key path is mistakenly identified as a list index by isdigit(), triggering current.extend([None] * (index - len(current) + 1)), attempting to allocate hundreds of millions of elements, instantly exhausting memory.
Root Cause Chain:
The key path contains a pure numeric field (such as group number 871684833)
→ isdigit() mistakenly identifies it as an array index
→ extend([None] * (871684833 - len(current) + 1))
→ attempts to allocate hundreds of millions of elements
→ memory exhausted → container OOM Kill (exit code -9)
Affected Version: 2.5.1 - 2.5.5
Fixed Version: 2.5.5
Repair Content:
- Always use a dictionary when pre-creating intermediate layers, no longer guess the container type based on whether the next segment is a number
- Only when the container itself is a list and the index is less than
STORAGE_MAX_LIST_INDEX(10000) do we handle it as an index, skipping large indexes safely - Change the recursive implementation to an iterative one, eliminating potential infinite recursion risks in the original code
- Add
STORAGE_MAX_LIST_INDEXconstant toCore/constants.py, centrally managing the index safety limit
Repair Date: 2026/07/10
Reproduction Steps:
# Writing a nested key path containing a large number field (such as QQ group number) triggers OOM
await sdk.storage.aset("groups.871684833.name", "某群")
# → Process memory surges instantly, killed by OOM
Regression Test: tests/unit/test_unit_storage.py adds 4 regression test cases
test_nested_key_numeric_segment_as_dict_key— precisely reproduces OOM scenariotest_nested_key_numeric_segment_multiple— multiple consecutive numeric fields all as dictionary keystest_nested_key_existing_list_index_set_within_limit— existing list index write within limittest_nested_key_list_index_safety_limit— safety limit verification for large indexes
Severity: 🔴 Severe
Type: Storage
[BUG-025] on_config_update callback not called by core route
Problem: on_config_update(old, new) callback is defined in the base class (BaseModule / BaseAdapter), but the core framework does not associate these events with configuration change events. As a result, when the configuration is modified through the configuration management panel, the callback is triggered, but when the config.toml file is manually edited or setConfig() is called, the on_config_update is not triggered.
Root Cause: ConfigManager emits config.set / config.updated lifecycle events when the configuration is changed, but there is no subscription logic to forward these events to each component's on_config_update method.
Root Cause Chain:
Core does not subscribe to config.set / config.updated
→ Configuration change events are not forwarded
→ on_config_update is not called
→ Manual editing file / code setConfig does not trigger hot update callback
Affected Version: All versions
Fixed Version: 2.6.2
Repair Content: ModuleManager / AdapterManager register config.set (covers code setConfig() path) and config.updated (covers manual editing file path) event subscriptions, match by configuration key prefix and call the corresponding component's on_config_update, passing type-safe configuration objects. Also fix _flush_config() not synchronizing _config_mtime after writing the file, avoiding the framework's own writing being mistakenly judged as external modification by the file monitoring task, which would repeatedly trigger config.updated.
Compatibility Note: The framework core now maintains configuration hot updates uniformly. Previously, the logic of triggering by the configuration management panel was removed, so the panel must be upgraded synchronously after upgrading the framework, otherwise duplicate triggers will occur (core + panel each trigger once). The on_config_update method signature and semantics remain unchanged, subclasses do not need modification.
Repair Date: 2026/07/23
Severity: 🟡 Medium
Type: Configuration System
[BUG-026] notice/request event reply target inference error
Problem: In group notification events (such as member joining a group group_member_increase), calling event.reply() sends the message to the private chat of the user who triggered the event, not to the group where the event occurred. The same applies to friend notification events, where the reply target may be confused.
Root Cause: infer_receive_type() directly returns the detail_type as the session type. For message events, this is correct (the detail_type values private/group are session types), but for notice/request events, the detail_type is a semantic subtype (such as group_member_increase, friend_increase), not a session type. Subsequent convert_to_send_type() and get_id_field() cannot find the value in the mapping table, so it falls back to the default "user" / "user_id", causing the reply target to be confused.
Root Cause Chain:
notice event detail_type="group_member_increase"
→ infer_receive_type() directly returns "group_member_increase"
→ convert_to_send_type("group_member_increase") not in mapping table → fallback "user"
→ get_id_field("group_member_increase") not in mapping table → fallback "user_id"
→ target_id = event["user_id"] ← new member's private chat (not group)
Affected Version: All versions
Fixed Version: 2.7.0-dev.3
Repair Content: infer_receive_type() adds a check—only return detail_type directly if it is a known session type (standard type or custom type); otherwise, infer the correct session type based on the ID field (group_id / channel_id / user_id etc.).
Regression Test: tests/unit/test_unit_session_type.py → TestNoticeRequestTypeInference (10 test cases)
Repair Date: 2026/07/29
Severity: 🟢 Minor
Type: Event System
[BUG-027] Route rate limit cleanup task uses fixed window causing long window rate limit rules to fail
Problem: When routing rate limit configuration is set to a long window rule (such as 100/hour, {"requests": 100, "window": 3600}), the rate limit is ineffective—practically behaving like 100/minute (up to about 6000 requests per hour), completely failing to provide the expected hourly protection.
Root Cause: _apply_rate_limit parses the actual window of each route (maximum 3600 seconds), and per-request checks also use this window; however, the background cleanup task _cleanup_expired_rate_limits uses a fixed constant DEFAULT_RATE_LIMIT_WINDOW_SECS (60 seconds) as the unified cleanup threshold for all routes. Thus, time stamps earlier than 60 seconds are cleared by the cleanup task, so the hour window never accumulates close to 100 records, and the rate limit is severely weakened.
Root Cause Chain:
_apply_rate_limit parses window=3600 (100/hour)
→ per-request checks use 3600s retention time (correct)
→ but _cleanup_expired_rate_limits uses fixed max_window=60s for cleanup
→ time stamps earlier than 60 seconds are cleared
→ the hour window only retains records from the last 1 minute
→ 100/hour effectively degrades to ~100/minute (relaxed by about 60 times)
Affected Version: 2.6.0-dev.0 - 2.7.0-dev.4
Fixed Version: 2.7.0-dev.5
Repair Content: Add _rate_limit_windows: dict[str, int] to record the actual window for each route; _apply_rate_limit writes the window when creating the entry for the first time; _cleanup_expired_rate_limits changes to clean up according to each key's own window (fallback to default value if missing); clean up deleted entries and stop() synchronize maintenance of the two dictionaries.
Repair Date: 2026/07/31
Regression Test: tests/unit/test_unit_router.py → TestRateLimit::test_cleanup_respects_per_route_window
Severity: 🔴 Severe
Type: Router
[BUG-029] Configuration listener task broadcasts incomplete TOML and silently swallows exceptions
Problem: When a user manually edits config.toml and saves it halfway (creating a temporary syntax error), the configuration listening background thread detects the mtime change, reloads the configuration, but fails to load and still broadcasts an empty configuration {} as config.updated, causing adapters/modules' on_config_update to receive an empty configuration, mistakenly assuming all configuration items have been cleared and reverting to default values. Additionally, the listener loop uses except Exception: pass to silently swallow all exceptions, making it impossible to troubleshoot watcher failures.
Root Cause: Two overlapping defects:
_load_configoverwritesself._cacheas{}when TOML syntax errors/permission errors occur, but the background listener thread_watch_loopand cache timeout path_check_cache_validityboth execute_emit_config_updated()unconditionally after calling_load_config(), broadcasting the "empty cache" generated by failed loading as a real change._watch_loopusesexcept Exception: passto not log any errors.
Root Cause Chain:
User saves halfway → TOML syntax error
→ _load_config() overwrites _cache = {}
→ _watch_loop unconditionally _emit_config_updated(new_config={})
→ Adapters/modules on_config_update receive empty config
→ Mistakenly assume configuration has been cleared, revert to default values
Affected Version: 2.6.2-dev.1 - 2.7.0-dev.4
Fixed Version: 2.7.0-dev.5
Repair Content:
_load_configis changed to returnbool; if TOML syntax errors/permission errors/other errors occur, the last valid cache is retained (not overwritten as{}), and diagnostic logs are recorded, returningFalse._watch_loopand_check_cache_validityonly emitconfig.updatedif_load_config()returnsTrue._watch_loop'sexcept Exceptionis changed to log at warning level (new i18n keycore.config.watcher_error, synchronized in five languages).
Repair Date: 2026/07/31
Regression Test: tests/unit/test_unit_config.py → test_malformed_toml_preserves_last_valid_cache, test_permission_denied_logs_clear_message (updated to verify retaining cache and returning False)
Severity: 🟡 Medium
Type: Configuration System
[BUG-030] Configuration watcher race condition causes setConfig delay write silently loses data
Problem: Multiple users report that after using config.setConfig(key, value) (default immediate=False), their module configuration is not written to config.toml, while other modules' configurations are normal. Setting immediate=True (force flush) can avoid this. Manifestation: Configuration written during runtime is lost after the next restart, but the configuration generated during startup remains.
Root Cause: Two overlapping defects:
- Logical defect:
_watch_loopunconditionally_dirty_keys.clear()when_check_file_change()returnsTrue, discarding all pending write keys. However,_check_file_change()only uses!=to compare mtime, and the framework's own_flush_configwriting also changes mtime—although_flush_configupdates_config_mtimeafter writing, the watcher thread may still observe mtime differences (and on coarse-grained file systems) between file writing and mtime assignment, mistakenly judging it as "external modification" and clearing all pending write keys. - Thread defect:
_watch_loopoperates_write_timer/_dirty_keyswithout holding_lock, creating data races withsetConfig(holding lock to write_dirty_keys),_schedule_write(holding lock to write_write_timer).
Root Cause Chain:
ModuleA setConfig(immediate=True) → flush writes to disk, mtime changes
→ User module setConfig(immediate=False) → enters _dirty_keys, flushes to disk after 5s
→ Watcher polling, _check_file_change observes mtime difference from its own previous write
→ _dirty_keys.clear() → User module's pending write keys are silently discarded
→ Configuration missing after restart
Affected Version: 2.6.0 - 2.7.0
Fixed Version: 2.7.1
Repair Content:
- Add
_last_self_write_mtimefield,_flush_configrecords it after writing;_check_file_changefirst compares this value when mtime changes, returnsFalseif matches, judging it as self-write _watch_loopholds_lockthroughout; retains_dirty_keysfor true external modifications (merge semantics), merges with external content on next flush (dirty keys take precedence), no longerclear()getConfig/_check_cache_validitypaths are unaffected (their reload does not clear dirty keys)
Repair Date: 2026/08/06
Regression Test: tests/unit/test_unit_config.py → test_self_write_not_detected_as_external, test_external_change_preserves_dirty_keys, test_flush_merges_dirty_with_external
Severity: 🔴 Severe
Type: Configuration System
[BUG-032] Reading immediately after writing configuration results in reading old values
Problem: config.setConfig() (default immediate=False delays writing for about 5 seconds) writes a dot-separated key, and immediately reads its parent/ancestor node (e.g., set_erispulse_section("scope.actions.MyModule", {...}) then calls get_erispulse_config()) returns the old value, the written sub-key "disappears" until the flush, affecting scenario such as scope configuration hot updates ("write-read-write" scenario) (2.8.0 test plugin /t_section case exposed).
Root Cause: setConfig stores dot-separated keys in the pending write queue _dirty_keys in a flat form, and only getConfig's exact key query hits the pending write queue; tree path queries (e.g., getConfig("ErisPulse.scope")) only go through the cache tree, not overlay pending write values—during the delay write period (_flush_config merges dirty keys into the cache and clears the queue), a read-you-write disconnect occurs.
Affected Version: 2.6.0 - 2.8.0-dev.1
Fixed Version: 2.8.0-dev.1
Repair Content: getConfig introduces pending write overlay semantics—① exact match in pending write key returns directly (original behavior unchanged); ② pending write key is the query key's ancestor → take the longest pending write ancestor and parse the remaining path in its value subtree; ③ pending write key is the query key's descendant → build an overlay subtree (_dirty_overlay) and deeply merge it with the cache subtree (_deep_merge, override priority, without modifying the original cache object). When there are no pending write keys, go through the original fast path, zero additional overhead.
Repair Date: 2026/09/04
Regression Test: tests/unit/test_unit_config.py → test_get_config_overlays_dirty_descendant, test_get_config_overlay_merges_with_cache_siblings, test_get_config_overlay_new_branch, test_get_config_dirty_ancestor_query, test_get_config_dirty_exact_key_still_wins
Severity: 🟡 Medium
Type: Configuration System
[BUG-033] wait_reply hanging reply is starved by high-priority handler
Problem: When a module calls wait_reply() to wait for a user reply, if that reply message is claimed by a higher-priority event handler (mark_processed()), the command dispatcher checks the _processed flag at the entrance and directly returns, so the reply matching _check_pending_reply hanging at the end of _handle_message never executes—the waiting party does not receive the reply and can only wait until timeout returns None. Typical triggering scenario: robots using high-priority message handlers (recording/auditing/blocking), all dialog interactions randomly fail.
Root Cause: The reply matching _check_pending_reply is placed at the end of _handle_message (only executed when command matching fails), while the _processed check is before it—the order of claim checking and reply matching is reversed. Interaction waiting is a framework-level session continuation mechanism, not a competitive handler, and should not be affected by other handlers' claim.
Affected Version: 2.2.0-dev.0 - 2.8.0-dev.1
Fixed Version: 2.8.0-dev.2
Repair Content: Move the reply matching judgment to the entrance of _handle_message (before _processed check, only for message events): first try to complete the pending conversation, if matched, the event is marked as processed, and the subsequent check naturally short-circuits; if not matched, continue the original command matching process. Also, delegate the judgment chain to a new interaction session manager (Core/Event/interaction.py), incidentally gaining session cancellation and permission review capabilities by affiliation.
Repair Date: 2026/09/08
Regression Test: tests/unit/test_unit_interaction.py (TestRegisterResolve match/not match/claim flag), tests/unit/test_unit_event.py (wait_reply full chain)
Severity: 🟡 Medium
Type: Event System / Command System
[BUG-034] persist=False runtime binding is silently overwritten by any subsequent configuration write
Problem: scope.set_module(..., persist=False) and other runtime bindings only modify memory self._data; but scope subscribes to config.set / config.updated events, and any code writing configuration (such as a module loading its own default configuration) triggers scope to rebuild the configuration tree from the configuration file, causing all previous runtime bindings to be silently lost (reverting to default allow), with no log prompts. Scenarios dependent on runtime bindings (Dashboard "runtime-only" switches, module runtime dynamic disabling) revert to behavior after unrelated module writes configuration.
Root Cause: Root cause chain: scope.set/delete(persist=False) only writes to memory (Core/scope.py) → any setConfig triggers config.set event → _on_config_updated unconditionally _load_config() → _apply_tree() replaces self._data = {...} as a whole → runtime bindings not in the configuration file are discarded.
Affected Version: 2.8.0-dev.1 - 2.8.0-dev.2
Fixed Version: 2.8.0-dev.2
Repair Content: Introduce a runtime override layer _runtime_overrides (with deletion sentinel): persist=False writes/Deletes are recorded in the override layer, _apply_tree() rebuilds the persistent layer after applying the new configuration, and runtime rules remain effective after any configuration write; persist=True writes/Deletes clear corresponding override records (user persistence semantics take precedence); config.set filters by event key precisely, config.updated compares new and old scope sections, and only rebuilds if the scope actually changes (with avoiding unrelated writes flushing LRU cache); add unregister_by_owner() for modules to clean up with the caller at unload. Core.Event.overrides has a similar issue with persist=False runtime overrides, which is simultaneously fixed with the override layer architecture.
Repair Date: 2026/09/09
Reproduction Steps: ① scope.set_module("testplat", blocked=["TestB"], persist=False) → judgment False; ② Any module executes config.setConfig("HelpModule", {...}) → triggers scope rebuild; ③ scope.is_allowed("testplat", None, "TestB") returns True (expected still False).
关联: Issue #432
Regression Test: tests/unit/test_unit_scope.py::TestRuntimeOverrideSurvival (irrelevant write survives/ tree rebuild replays/ deletion sentinel/ persistent clear/ precise invalidation/ owner cleanup)
Severity: 🟡 Medium Type: Configuration System / Runtime
[BUG-035] Configuration panel select options and dict fields render as [object Object]
Problem: In the WebUI configuration panel, select field options display [object Object] (such as dynamically generated color style options); dict fields without declared control types (such as stalker_mode, knowledge_base, etc.) display [object Object] in text boxes, making it impossible to view and edit normally.
Root Cause: Two independent defects: ① The framework i18n resolver _resolve_i18n_text only restores dictionaries with i18n keys, and option labels are dictionaries with only default (no i18n key for dynamic text) which are passed through as is, resulting in [object Object] after frontend esc(label) string coercion; ② The Dashboard rendering branch only handles JSON textarea for array types, dict values fall into the plain text input branch and are converted by String(); additionally, if a module mistakenly declares _schema_meta as a regular dataclass field (missing ClassVar annotation), it is treated as a configuration field and enters the schema, exacerbating confusion.
Affected Version: 2.7.0 - 2.8.0-dev.2
Fixed Version: 2.8.0-dev.2
Repair Content: ① _resolve_i18n_text supports dictionaries with only default as text; ② Framework schema/template/default value/filling/validation five places exclude fields with underscore prefix (mistaken declaration harmless); ③ Dashboard select option label object fallback parsing (priority default), dict/table fields rendered as JSON textarea (save path by tp=object JSON.parse back, complete round trip).
Repair Date: 2026/09/09
Regression Test: tests/unit/test_unit_config.py::TestResolveI18nDefaultOnlyDict, TestSchemaUnderscoreFieldExclusion
Severity: 🟡 Medium Type: Configuration System
[BUG-036] Multi-instance shared configuration directory causes occasional configuration write failures (ENOENT)
Problem: In Docker deployment scenarios (multiple containers mounting the same host configuration directory), the log occasionally shows two consecutive lines of Failed to write configuration file ... [Errno 2] No such file or directory: '...config.toml.tmp' -> '...config.toml'. This configuration write is discarded (the old configuration is fully retained, no configuration loss is observed), functionality is unaffected, but the alert repeatedly interferes with troubleshooting, and pending configuration items must wait for the next write to be written to disk.
Root Cause: Root cause chain: _flush_config / setConfigTemplate uses a fixed-name temporary file config.toml.tmp to carry new content, write() does not fsync before rename() directly. Two ErisPulse instances sharing the same configuration directory, B instance open("w") may truncate A instance's ongoing temporary file → A rename() when the target has been taken or content truncated by B, reports ENOENT (i.e., the user's log of two consecutive errors). In reported cases, only write failure alerts were observed (the old configuration was retained); if the timing overlap is more extreme, rename may output empty/half-written config.toml (a potential risk, not yet exploded in real environments). In single-instance scenarios, ext4's delayed allocation also has a "rename metadata before data block is written" crash window (SIGKILL / power failure). _file_lock is a process-level threading.RLock, which has no constraint on cross-process/cross-container writes.
Affected Version: 2.2.0-dev.0 - 2.8.0
Fixed Version: 2.8.1
Repair Content: Converge all configuration writes to _atomic_write_text(): generate a process-unique temporary file (eliminate fixed-name competition, multiple instances degrade to last-writer-wins, no longer ENOENT) → write and flush + fsync force data to disk (eliminate "rename effective, data not written" window) → os.replace atomically replace target (POSIX/Windows are atomic, at any moment the disk has either complete old content or complete new content); additionally fsync the configuration directory on POSIX. Three write points (_flush_config, setConfigTemplate, root directory configuration migration) all switch; exception path cleanup logic is restructured with the unique temporary file name. Add multi-instance detection: during startup, an advisory lock (flock on POSIX / msvcrt.locking on Windows) exclusively holds the configuration directory lock file .erispulse_config.lock, if occupied, output an i18n warning (does not block startup), and the lock is automatically released by the OS when the process exits, with no ghost lock.
Repair Date: 2026/09/13
Reproduction Steps: ① Two containers mount the same host config/ directory and run ErisPulse simultaneously; ② Trigger configuration write (such as module registration default configuration) in any instance; ③ Observe log errors of ENOENT write failure, this write is discarded (old configuration retained).
关联: User report (1Panel container ×2)
Regression Test: tests/unit/test_unit_config_atomic_write.py (complete content write / no temporary file residue / write failure preserves original file / dual-instance concurrent write always valid file / lock file creation / multi-instance warning / migration atomic write)
Severity: 🟢 Minor
Type: Configuration System
[BUG-037] Commands with space-separated names cannot be triggered after registration
Problem: After registering a sub-command with a space-separated command name (such as @command("admin add")), the command can be registered normally and appears in the help list, but when the user sends /admin add, the robot never responds—input is matched as admin plus parameter ["add"] by the parent token command; if the parent token is also unregistered, there is no response at all. Only when using dot-separated naming (such as admin.reload, as a single token) can it be avoided.
Root Cause: Root cause chain: CommandHandler.__call__ stores any command name (including space-separated forms) as a key in a flattened self.commands dictionary → during distribution, _try_execute_command only matches the first message token (cmd_name = parts[0]) → multi-token command names as dictionary keys are never found. The registration and matching stages have inconsistent assumptions about the command name space, and there is no registration-time warning (silent failure).
Affected Version: Since the introduction of the command system - 2.8.0
Fixed Version: 2.8.1
Repair Content: The matching layer changes to longest prefix matching: try matching from the longest candidate (" ".join(parts[:n]), n capped by the maximum token count of registered command names/aliases _max_name_tokens) and degrade stepwise, matching triggers execution with the remaining tokens as parameters, downstream scope/ACL/overriding/master/permission chain naturally applies to the full command name. Accompanying semantics: when parent and child coexist, unregistered sub-command input falls back to the parent command (unchanged historical behavior); if a sub-command does not declare permission, it inherits the most recently declared ancestor command's permission (protecting the parent command protects all its sub-commands); only registering single-token commands matches on the first round, distribution overhead is consistent with the original. unregister / unregister_by_owner / full cleanup synchronously maintains the token count cache.
Repair Date: 2026/09/13
Reproduction Steps: ① Module registers @command("admin add"); ② Sends /admin add x; ③ Before repair, there is no response (or is caught by the unregistered /admin as a parameter), after repair admin add triggers and get_command_args() is ["x"].
Regression Test: tests/unit/test_unit_command_subcommand.py (longest prefix matching / only sub-command triggers / three-level nesting / case-sensitive two modes / single and multi-token aliases / event payload full name / lifecycle hooks full name / permission inheritance six examples / ACL glob full name / master / unregistration fallback and cache recalculation)
Severity: 🟡 Medium
Type: Event System
[BUG-038] persist=False runtime binding is written to disk along with persistent write, module unloads from disk "revive"
Problem: Runtime bindings written with scope.set(path, value, persist=False) (documented promise not to write to disk, invalid after process restart, cleared when module unloads) are written to disk along with any unrelated persist=True write (default value, such as module set_module / WebUI save configuration); after the module unloads and unregisters the runtime binding, it is still "revived" from disk after the next configuration reload and continues to be effective after process restart—contrary to the "runtime binding does not write to disk" semantic contract, and extremely difficult to troubleshoot.
Root Cause: Root cause chain: ScopeManager.set() first writes the value to the memory configuration tree _data (runtime bindings are also directly written to _data) → the persistent branch makes a deep copy snapshot of the entire _data tree and submits it to update_erispulse_config for differential writing → the snapshot includes the values of persist=False bindings. delete(persist=True) is the same source: the live reference (parent node) is given to the delayed write dirty queue, and during the delayed write period, the live reference is in the dirty queue, causing a cross-write contamination window.
Affected Version: 2.8.0-dev.2 - 2.9.0-dev.0
Fixed Version: 2.9.0-dev.1
Repair Content: Introduce a persistent baseline _persisted_tree (the configuration tree is refreshed from disk-loaded, validated tree as the mirror of disk truth): set(persist=True) only applies the current change to the baseline and submits the difference for writing, runtime bindings never enter the persistent content; "write-then-read" is changed to restore the memory final state snapshot directly, no longer through _apply_tree reconstruction to pollute the baseline. delete(persist=True) is the same口径: apply the deletion to the baseline and give the deep copy of the baseline subtree to the persistent layer; set_action whole replacement semantics first deletes according to the same口径 then writes, old rule keys are not left in the persistent content.
Repair Date: 2026/09/27
Reproduction Steps: ① Module executes scope.set("bots.p.debug_mode", {"blocked": ["X"]}, persist=False); ② Trigger any persistent write (such as another module calling set_module); ③ Open config/config.toml, debug_mode has been written (before repair); ④ Unload the module (runtime binding is unregistered) then trigger configuration reload, scope.get("bots.p.debug_mode") still returns the bound value.
Regression Test: tests/unit/test_unit_scope.py::TestPersistBaseline (runtime bindings do not write to disk with unrelated persistent writes / unregistered after unload does not revive on configuration reload / delete submits baseline subtree without runtime sibling keys / set_action replacement leaves no residual keys / cache_size configuration takes effect)
Severity: 🔴 Severe
Type: Configuration System
[BUG-039] Inconsistent read-write when section write and dot write coexist (dot overwrite lost)
Problem: During the delay write window, when both section write (setConfig("Mod", {...}), such as BaseModule.cfg writing back) and dot write (setConfig("Mod.key", v), such as configuration hot update / test tool overwrite) coexist, there are two variants: Variant A (read path) — getConfig("Mod") section read returns the old section pending write snapshot, not seeing the later dot overwrite (dot read path is normal); Variant B (write path, heavier) — dot write before, section write after (test tool injection overwrite → module self.cfg = ... writing back common timing) — flush applies dirty keys in insertion order, section write completely replaces the section, dot overwrite is permanently lost on disk. Production environment "configuration hot update + module runtime write back" combination can trigger, unrelated to test environment (ErisPulse-DailyCard test feedback exposed).
Root Cause: getConfig first segment (exact match in _dirty_keys) returns early return self._dirty_keys[key], completely bypassing the fourth segment _dirty_overlay descendant overlay; _flush_config applies dirty keys in insertion order, section write happens after dot write, so dot values are overwritten. Also: getConfig returning (exact match / ancestor subtree / overlay merge) in the first segment returns (reference to dirty queue object), modifying the returned dict in place directly alters the pending write state.
Affected Version: 2.6.0 - 2.9.0-dev.1
Fixed Version: 2.9.0-dev.1
Repair Content: In the dirty window, unify specificity priority semantics — dot (more specific) pending values take precedence over section (wider) pending values, read path and flush path same口径: ① getConfig after exact match in pending key still overlays _dirty_overlay descendant pending values (not dict values return overlay subtree, same口径 as ③+④ scalar edges); ② _flush_config applies dirty keys in path depth order (section/ancestor first, dot/descendant later, stable sort maintains same depth write order). Cost: In the same dirty window (about 5 seconds), section write back cannot overwrite still pending dot write — read path fixed, read-modify-write naturally carries dot values, actual impact surface is extremely small. ③ Involved pending values of getConfig return (exact match / ancestor subtree / overlay merge) changed to copy.deepcopy isolated copy, no longer leak dirty queue internal reference; no dirty key fast path and pure cache read behavior unchanged.
Repair Date: 2026/09/28
Reproduction Steps: ① setConfig("FB.y", 1); ② setConfig("FB", {"z": 2}); ③ force_save() — before repair, disk only has [FB] z = 2 (y lost), after repair {"y": 1, "z": 2}; Variant A: steps ①② without flush directly getConfig("FB") — before repair does not contain y, after repair visible.
Regression Test: tests/unit/test_unit_config.py::TestSectionAndDottedDirtyConsistency (Variant A two insertion orders / Variant B flush and order independent / three-layer mixed survival / isolated copy / ancestor subtree isolation)
Severity: 🟡 Medium
Type: Configuration System