Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
sjaakola
MDEV-41169 Memory leak in wsrep_sst_prepare

Error handling in sst_prepare_mysqldump() could overwrite
the error code from malloc/sprintf with status received later from
mysql_thread_create() call. This let the function report success
after an allocation failure, without setting *addr_out,
breaking the caller's ability to correctly free the allocated
address buffer.

This commit fixes the error handling in sst_prepare_mysqldump so
that each error path frees its own allocation. Thread create uses
separate status variable.
Kristian Nielsen
Binlog-in-engine: fix couple typos

Signed-off-by: Kristian Nielsen <[email protected]>
Arcadiy Ivanov
MDEV-36832 UPDATE...IN(SELECT) acquires spurious gap locks on FK index

`UPDATE ticket SET is_valid=0 WHERE ticket_id IN
(SELECT ticket_id FROM ticket WHERE booking_id=1)` acquires S next-key
and gap locks on the FK secondary index `fk_ticket_2_booking`, causing
deadlocks between concurrent transactions operating on non-overlapping
`booking_id` values. The expected behavior — matching
`UPDATE ... WHERE ticket_id IN (1,2)` (explicit PK list) — is to only
lock the PRIMARY key records being updated.

**Root cause**: the optimizer converts the `IN(SELECT)` to a semi-join,
which creates two table references for the same table. The SQL layer
then processes the statement as `SQLCOM_UPDATE_MULTI`. In
`ha_innobase::store_lock()`, the read-only (subquery) table handle
receives `TL_READ`, but the existing `LOCK_NONE` optimization for DML
read tables was gated on `isolation_level <= READ_COMMITTED` and did
not include `SQLCOM_UPDATE_MULTI` at all. In `REPEATABLE READ` (the
default), the subquery's secondary index scan therefore acquired
`LOCK_S` next-key locks, including gap locks spanning into adjacent key
ranges.

**Fix**: add `SQLCOM_UPDATE_MULTI` to the `LOCK_NONE` (consistent read)
condition in `store_lock()`:

1. For `TL_READ` (binlog off or ROW format): use consistent read at
  all isolation levels below `SERIALIZABLE`. This is safe because:
  - The MVCC snapshot is stable within a transaction in
    `REPEATABLE READ`
  - Rows being modified still acquire X locks via the write-side
    table handle (`F_WRLCK` / `LOCK_X`)

  `SERIALIZABLE` is excluded: the `::external_lock()` upgrade of
  `LOCK_NONE` to `LOCK_S` applies only to non-autocommit
  transactions, so an autocommit multi-table UPDATE would otherwise
  read its read-side tables via a consistent read at `SERIALIZABLE`,
  weakening its guarantees.
2. For `TL_READ_NO_INSERT` (statement-based binlog): extend the
  existing `READ COMMITTED` optimization to also cover
  `SQLCOM_UPDATE_MULTI`

The fix is deliberately limited to `SQLCOM_UPDATE_MULTI` (not
`SQLCOM_UPDATE`) because single-table UPDATE with scalar subqueries on
other tables traditionally uses S locks to block concurrent reads of the
subquery table, and broadening the change there would alter observable
behavior for existing applications.

**Lock count before/after** (semi-join UPDATE on `booking_id=1`):
- Before: 5 lock structs, 5 row locks (IS + S×3 on FK index + IX +
  X×2 on PK)
- After: 2 lock structs, 2 row locks (IX + X×2 on PK) — identical to
  explicit `WHERE ticket_id IN (1,2)`

**Tests**:

`update_subquery_fk_gap_lock.test` — verifies that concurrent
`UPDATE...IN(SELECT)` operations on non-overlapping `booking_id` values
do not block each other. Two connections each update their own
`booking_id`'s tickets via semi-join subquery, then insert new tickets
for their `booking_id`. Both operations must succeed without lock wait
timeout, confirming that the spurious FK index gap locks are eliminated.

`update_subquery_fk_gap_lock_isolation.test` — ACID isolation regression
tests (requires ROW binlog format) verifying that replacing S locks with
consistent read (`LOCK_NONE`) on the semi-join read table does not cause
isolation violations at `REPEATABLE READ`:

1. **Phantom insert**: T1 updates `booking_id=1` tickets via semi-join,
  T2 concurrently inserts a new ticket with `booking_id=1`. T2 must
  not block (no S gap locks). T1's MVCC snapshot does not see the new
  row, so T1 updates only the 2 original tickets. Correct under
  snapshot isolation.
2. **Concurrent UPDATE on filter column**: T1 updates `booking_id=1`
  tickets, T2 changes `ticket_id=2`'s `booking_id` from 1 to 2. Both
  need X lock on `ticket_id=2`'s PK record — serialized via PK X
  lock alone, proving FK index S locks are unnecessary.
3. **Concurrent DELETE**: T1 updates `booking_id=1` tickets, T2 deletes
  `ticket_id=1`. Both need PK X lock — serialized, no isolation
  violation.
4. **SERIALIZABLE S locks**: At `SERIALIZABLE`, the read-side table
  keeps `LOCK_S`, so T2's INSERT into the subquery range must block
  with `ER_LOCK_WAIT_TIMEOUT`. Confirms the fix does not weaken
  `SERIALIZABLE` guarantees.
5. **Multi-table UPDATE on different tables**: T1 does
  `UPDATE t_write JOIN t_read ... SET t_write.val = t_read.val`, T2
  concurrently modifies `t_read`. T2 must not block (T1 holds no S
  locks on `t_read`). T1's snapshot values are used. Correct under
  snapshot isolation.
6. **MVCC snapshot consistency**: T1 establishes snapshot, T2 modifies
  `t_read`, T1 does multi-table UPDATE using snapshot then re-reads.
  Re-read must return the same values the UPDATE used, confirming
  MVCC snapshot stability at `REPEATABLE READ`.
7. **SERIALIZABLE with autocommit**: the `::external_lock()` upgrade
  of `LOCK_NONE` to `LOCK_S` does not apply to autocommit statements,
  so `store_lock()` itself must not choose `LOCK_NONE` at
  `SERIALIZABLE`. An autocommit `UPDATE...IN(SELECT)` stalled
  mid-scan on a PK X lock must hold S next-key locks on the FK index
  entries it has already scanned: an INSERT whose FK index entry
  falls inside that range must block with `ER_LOCK_WAIT_TIMEOUT`.

`mdev-14846.test` depends on the multi-table UPDATEs' read-only
tables being read with S locks to produce the deadlock whose error
handling it verifies. Both transactions taking part in that deadlock
now run at `SERIALIZABLE`, where the read-only tables keep their S
locks, preserving the original deadlock scenario and its
`ER_LOCK_DEADLOCK` expectation unchanged. Tying this to the binary
log format instead would not work everywhere: the embedded server
has no binary log, so its read side always gets `TL_READ`.
Alexey Yurchenko
MDEV-38920 MTR tests for Galera-side fixes

MDEV-38920-evs-config-warn checks that there is a warning about bad
configuration values and they are not accepted.
MDEV-38920-install-timer-expired reproduces 'install timer expired'
situation.
Both tests require fixed Galera library to pass (4.29) hence will be
skipped when run with older versions.
Andrzej Jarzabek
MDEV-37521 Attempt to compare iterators from different sequences in range_set::add_range

fil_space_t::freed_ranges is a range_set guarded by
freed_range_mutex, and every mutator takes that lock except one:
the branch of mtr_t::commit() that frees pages for an unlogged
mini-transaction on the temporary tablespace (the common case,
since temp-space pages never set m_modifications). That path
called fil_space_t::add_free_range() directly, racing with any
other locked mutator of the same std::set - most notably
fil_space_t::flush_freed(), which the page cleaner invokes for
every tablespace while resizing the buffer pool. The race
corrupts the range_set's underlying tree and crashes the server.

Take freed_range_mutex around the loop, matching every other
caller (process_freed_pages(), flush_freed(),
clear_freed_ranges()).

Added innodb.temp_truncate_freed_race, which reproduces the
crash with a configurable probability by racing concurrent
temporary-table churn against a continuously resizing buffer
pool across several rounds against the same server.
Vladislav Vaintroub
MDEV-40967 PROXY protocol host check sent in clear text mid-SSL handshake

Defer the host-privileged/host-blocked check for a PROXY-header-derived
address until after the client's SSL handshake completes, instead of
sending it immediately in clear text.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Thirunarayanan Balathandayuthapani
MDEV-40319 Instant ALTER TABLE rollback corrupts virtual column

Problem:
=======
ha_innobase_inplace_ctx::~ha_innobase_inplace_ctx() runs, whenever
ctx->instant_table is set and frees the old_v_cols exist also.
old_v_cols and old_n_v_cols are captured in the constructor as
prebuilt_arg->table->v_cols and n_v_cols, i.e. an alias of the live
table's own virtual columns, not a copy. By the time this destructor
runs, old_table->v_cols is either still that same array. If the
failure happened before ctx->instant_column() ever ran like during
prepare_inplace_alter_table_dict() or failure happened during the
commit phase innobase_instant_try(). In both cases old_v_cols is
the table's current, live v_cols array, so this loop destructs
dict_v_col_t objects that are still in use.

Solution:
========
ha_innobase_inplace_ctx::~ha_innobase_inplace_ctx(): Destruct
instant_table->v_cols[], not old_v_cols[]. instant_table is the
independently allocated dict_table_t that prepare_instant() built;
It owns its own v_cols array, whose dict_v_col_t::v_indexes must
be destructed before dict_mem_table_free() reclaims instant_table's memory.
Oleksandr Byelkin
MDEV-41258 Crash in Item_func_nextval on prepared re-execution

Sequence prelocking entries were matched across re-executions by
db.str/table_name.str pointer identity, which aliases memory owned by
the TABLE's own mem_root. Once the owning table was closed and reopened
(FLUSH TABLE, table-cache eviction), those pointers went stale and no
longer matched, leaving a dangling duplicate entry that still got
opened and crashed.
Fix: match entries by owning TABLE_LIST and position instead, and
deep-copy db/table_name onto the statement's own arena once, at entry
creation.

The relink was gated the same way as trigger/routine prelocking
discovery (has_prelocking_list), which stays true forever once a
statement's prelocking set includes a trigger, so it silently stopped
running after the first execution for any such statement.
Fix: moved it into its own function, called unconditionally per table,
scoped to DML statements only (not ALTER or INSERT DELAYED).

A stale entry's linked_table pointer could remain set when its owning
table's own open was skipped in a given execution, or when a
mid-statement open_tables() backoff (close_tables_for_reopen()) closed
the owner without a subsequent relink, so a later open_table() backlink
write could land on already-freed memory.
Fix: clear it in TABLE_LIST::reinit_before_use() and in
close_tables_for_reopen().

set_parameters() overwrote lex->default_used every execution from that
execution's own bound parameters only, losing track of a literal
DEFAULT written in the prepared statement text itself.
Fix: also record default_used at prepare time and OR it in.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Kristian Nielsen
Fix hang on master when disabling semi-sync

There is a global variable global_ack_signal_fd used to signal the receiver
thread to wake up when disabling semi-sync. This variable was cleared to -1
in the Ack_listener destructor, which ran at the end of the
Ack_receiver::run() function without any locking. If the thread was delayed
at that point, it could end up overwriting the new value set by a new
receiver thread. This would leave the server in a state with an invalid
global_ack_signal_fd and could cause a subsequent disable of semisync to
fail due to the wakeup not arriving at the receiver thread.

This was seen as a sporadic failure of the test case
rpl.rpl_semi_sync_cond_var_per_thd.

Fix by not modifying the global in constructor/destructor; instead set and
clear the global explicitly, allowing to clear the fd with proper locking
while the mutex is still being held.

Also fix a missing pthread_join(), which would leak thread descriptors and
allow to start a new receiver thread before the old one shut down fully.

(Either of these two changes fix the hang bug).

Signed-off-by: Kristian Nielsen <[email protected]>
FebArch
MDEV-40997: Remove unused mysqlbinlog_have_debug.inc

Verified via git grep that this file has zero references anywhere
in the codebase on the 10.11 branch.

Ran mysql-test-run.pl --suite=binlog: 96.97% pass rate, with one
pre-existing unrelated failure (binlog.binlog_mysqlbinlog-cp932) -
a known Windows console/CP932 encoding environment issue, also
observed identically on main.
Khaled Riyad
MDEV-40551 Copy/Paste friendly output format for MariaDB Command Line Client

Copy/paste friendly output was only reachable by starting the client with
--silent --skip-column-names, which cannot be done from a running
interactive session.

Add \S, a statement terminator which prints the result of one statement in
the tab separated format without column names.

com_silent() sets output_plain, opt_silent and column_names around
com_go(), then restores them, the same way com_ego() handles vertical.
output_plain selects print_tab_data() ahead of the vertical and table
branches, so \S gives the same output whether the session was started
plainly or with --table, --vertical or --silent. --html and --xml still
win, matching \G.
Thirunarayanan Balathandayuthapani
MDEV-38942  i_s_dict_fill_sys_tables() aborts when reading INNODB_SYS_TABLES after innodb_force_recovery

Problem:
========
  A query on INFORMATION_SCHEMA.INNODB_SYS_TABLES crashes when
SYS_TABLES contains a record that was inserted by a transaction
which has not been committed. This can happen after a
crash while a CREATE TABLE was in progress, if the server is
restarted with innodb_force_recovery=4 or greater,
because trx_rollback_recovered() is then skipped and the
recovered transaction remains ACTIVE.

dict_sys_tables_rec_read() returns READ_NOT_FOUND for such a
record, and dict_load_table_low() returns that as success with no
error message and setting *table to nullptr.
i_s_sys_tables_fill_table() checks only the error
message and passes the nullptr table to i_s_dict_fill_sys_tables(),
which dereferences it.

Solution:
========
i_s_sys_tables_fill_table(): Skip the SYS_TABLES record when
dict_load_table_low() reports success but returns no table, because
such a record is not visible.
Marko Mäkelä
MDEV-40410: Tight innodb_buffer_pool_size_max on ThreadSanitizer

bur_pool_t::size_in_bytes_max_default: Define as 0 also on
ThreadSanitizer. The symbol __SANITIZE_THREAD__ is predefined
starting with Clang 22 or GCC 7 when building with -fsanitize=thread.
Rucha Deodhar
MDEV-39049: Memory corruption & crash in check_key_in_list upon using
JSON_KEYS after modifying character set name/collation

Analysis:
Buffer overflow crashes and empty key duplication in check_key_in_list.

Fix:
Checking result buffer validity prevents segmentation faults on empty
strings and malformed inputs while preserving correct key matching behavior.
Marko Mäkelä
MDEV-41152: Fix FILE_CREATE recovery

fil_name_process(): Treat FILE_CREATE in the same way as FILE_MODIFY
that led to a FIL_LOAD_DEFER return. Remove the parameter lsn,
and return file_name_t& in which the caller may assign create_lsn
when processing a FILE_CREATE record.

deferred_spaces.reinit_all(): Never create anything for deleted
tablespaces. Doing so could cause a legitimate file to be deleted
if files are being deleted and re-created with the same name.

deferred_space.create(): Remove some duplicated code. Missing
tablespace files will be created in fil_node_open_file_low()
starting with
commit 759e3523e3d832b174cf0a612704da38b2557b40 (MDEV-38026).

deferred_spaces::item::lsn: Remove. Starting with
commit 37d8577aee3bd87b5b04464144d064063b169039 (MDEV-40728)
each FILE_ record is parsed only once.

recv_sys_t::parse_store_if_exists(): Tell the caller to skip
tablespaces for which both FILE_CREATE and FILE_DELETE was parsed.
This improves performance, not correctness.

recv_validate_tablespace(): Avoid duplicated tablespace lookup
and remove a redundant deferred_spaces.add(); fil_name_process
already keeps deferred_spaces in sync with recv_spaces.

fil_space_t::rename(): If !log, assert !replace and that
the target file name does not exist.

os_file_rename_func(): Do not check that the target path
does not exist. This is already checked by every caller.
This fixes a debug assertion failure that could otherwise
occur when recovering from a crash in
fil_space_t::rename() between the write of the FILE_RENAME
and the actual rename.

Reviewed by: Thirunarayanan Balathandayuthapani
Tested by: Saahil Alam
Alexey Yurchenko
MGL-299 Regression in galera_sst_rsync_encrypt_with_key MTR test

Commit b68e29a9c64 explicitly disabled use of SSL encryption in SST
by setting ssl-mode=DISABLED in the top configuration files.
This test is a backward compatibility test so it relies on the
deduction of ssl-mode from the presence of tkey and tcert params
in [sst] section. Unset ssl-mode in config to allow to derive it
from the presence of tkey and tcert.
Monty
Fixed that translog_walk_filenames() in Aria properly recognized aria
log filenames.
Aleksey Midenkov
MDEV-40821 SIGSEGV in Window_funcs_sort::setup

Window_funcs_sort failed accessing win_func->window_spec as window
spec was not defined in the query. Usually this is checked by fixing
item win_func, but it was not done by setup_conds.

Actually, earlier check by vers_setup_conds() must fail on DELETE
HISTORY from non-versioned table, but TABLE_LIST for t1 has no
versioning conditions.

The cause was the parser assigning versioning conditions to wrong
table pointed by last_table() which was already switched to another
table from SYSTEM_TIME expression (cs1).

The fix assigns versioning conditions to the correct table stored to
correspondent_table by delete_single_table branch of the parser.

Same fix applied to delete_single_table_for_period (period_conditions),
with tests added for both cases.

Also DBUG_ASSERT in TABLE::vers_switch_partition() checked
delete_history on last_table(), which is also wrong for this case.
Checking the query_tables is correct, as delete_single_table puts
the DELETE target first there, and tables referenced later in the
statement are appended after it.
Vladislav Vaintroub
MDEV-41080 startup code on Windows, remove checks for existing service

They were not necessary, just try to run as service, and fallback to
command line.

Add some diagnostics - unexpected errors from StartServiceCtrlDispatcher
and RegisterServiceCtrlHandler are now reported to Windows event log.

Also use authoritative service name, returned as first argument
in svc_main by service control manager.
Raghunandan Bhat
MDEV-41193: ASAN heap-buffer-overflow in ha_connect::CheckCond after select from Connect table

Problem:
  When CONNECT engine pushes a WHERE clause down to an external table,
  it writes the filter into the work area, without checking how much
  space is left. A large string literal in the WHERE clause can overflow
  the work area allocated by the engine. For ex: if connect_work_size is
  set to 4MB and the string literal in the WHERE clause is larger than
  4MB, it can grow past the allocated work area.

Fix:
  Track the space left in the work area and check it before writing. If
  the filter doesn't fit, drop it instead of writing past the buffer.
Marko Mäkelä
MDEV-41166 --backup --innodb-log-checkpoint-now may copy too much

xtrabackup_backup_func(): Request for a checkpoint synchronously
so that recv_sys.find_checkpoint() will observe the effect.

Reviewed by: Thirunarayanan Balathandayuthapani
Aleksey Midenkov
MDEV-41157 Follow-up test fix: strip version-dependent charset result

Follow-up to 6c391b92055.

SHOW CREATE DATABASE always appends "/*!40100 DEFAULT CHARACTER SET
... */", but the default charset differs by version (latin1 on 11.4,
utf8mb4 on 11.8+). Strip it instead of matching a specific value,
since it is not the subject for the fix.

Also switch to evalp: same execution, but logs "$var" instead of the
substituted value, so the replace_regex/disable_query_log hacks go away.
Marko Mäkelä
MDEV-33966: buf_page_make_young() is a contention point

The buf_pool.LRU list needs to reasonably accurately reflect recently
accessed blocks, so that they will not be evicted prematurely.
Because the list is protected by buf_pool.mutex, it is not a good idea
to maintain the position on every page access.

Instead of maintaining the LRU position on each access, we will
decide on each access whether the block qualifies for promotion, and
record the decision in a "promote" flag in the block descriptor. The
buf_flush_page_cleaner() thread as well as some traversal of the
buf_pool.LRU list, all of which hold buf_pool.mutex anyway, will move
the flagged blocks to the "recently used" end of buf_pool.LRU.

A block qualifies only if it is accessed again at least
innodb_old_blocks_time after the reference point, which is its first
access and, after that, its most recent promotion. The age is measured
in whole seconds and checked on the access, not when a sweep reaches
the block, so that a block that a table scan accessed in one burst
will not be promoted, however long it stays in buf_pool.LRU_old. After
a promotion, the age is measured from the time the sweep moved the
block, which can be later than the access that qualified it. The rule
applies at any position in buf_pool.LRU, and the flag stays set until
a sweep moves the block. Thus, a block that qualified outside
buf_pool.LRU_old keeps the flag until a sweep reaches it, which is
usually after it has moved into buf_pool.LRU_old; before, such a
block was moved on any access when freed_page_clock showed that it was
no longer close to the "recently used" end.

buf_page_make_young_if_needed(), buf_page_make_young(),
buf_page_peek_if_too_old(), buf_page_peek_if_young(),
btr_cur_nonleaf_make_young(), buf_page_t::set_accessed(): Replaced by
buf_page_t::touch(), buf_page_t::touch_no_stamp() and
buf_page_t::make_young_if_needed().

buf_pool_t::freed_page_clock, buf_page_t::freed_page_clock: Remove.
This is no longer meaningful in the revised design.
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).FREE_PAGE_CLOCK now
always reports 0. Before the first eviction, buf_LRU_stat_update() now
records statistics intervals, and buf_LRU_evict_from_unzip_LRU() uses
its formula instead of assuming a disk-bound workload.

page_zip_des_t::state: An atomic 16-bit field that will include
the PROMOTE and OLD flags that would more logically belong
to buf_page_t. We maintain them here (along with some
ROW_FORMAT=COMPRESSED specific state that is protected by page
latches) in order to avoid race conditions and unnecessary
memory overhead. m_end moves out to its own field, shrinking the
packed word to 16 bits; n_blobs narrows from 12 bits to 10 to make
room for PROMOTE and OLD, still comfortably above the 744-column
maximum.

The helpers that set or clear one bit of page_zip_des_t::state always
perform the atomic read-modify-write, which is a full barrier on x86.
A caller that can find the bit already in the target state checks it
first: touch_no_stamp() for PROMOTE, buf_LRU_add_block() for OLD, and
page_zip_write_rec() for NONEMPTY. page_zip_des_t::copy_from() and
set_n_blobs_and_empty() skip the fetch_add() when the delta is 0.
page_zip_decompress_clust() counts the BLOB pointers locally and
invokes add_n_blobs() once.

page_zip_des_t::copy_from(): Replaces the copy constructor that
page_zip_copy_recs() invoked. It copies m_end, NONEMPTY and n_blobs
from the source page, but preserves the PROMOTE and OLD flags of the
destination block, because they describe the position of that block
in buf_pool.LRU. Copying OLD from the source would make
buf_pool.LRU_old_len inconsistent with the list. Like
set_n_blobs_and_empty(), it updates the state with a single
fetch_add(), which preserves any concurrent change of PROMOTE or OLD.

buf_page_t::init(): Takes the ROW_FORMAT=COMPRESSED shift size
(ssize) directly and clears the zip descriptor state itself, so
callers (buf_block_t::initialise(), buf_page_init_for_read()) no
longer need a separate page_zip_des_init()/page_zip_set_size() call.

buf_page_t::invalidate(): Replaces buf_block_modify_clock_inc().
Instead of maintaining a 64-bit counter, we will maintain one
comprising 32+16=48 bits, in modify_clock_low,modify_clock_high.
Worst case there will be exactly n<<48 calls to
buf_page_t::invalidate() before some operation such as
btr_pcur_t::restore_position() is executed. Such a count should
be extremely unlikely but not completely impossible. It is worth
noting that the DB_TRX_ID is only 48 bits, and each transaction
start and commit/rollback will consume an identifier.

buf_page_t::modify_clock(): Replaces the read access of
buf_block_t::modify_clock. Assert that the caller is holding
a page latch. Note: because invalidate() and modify_clock() are
protected with buf_pool.mutex or the buf_page_t::lock, there can
be no issue with regard to the atomicity of accessing the 48-bit field.

buf_page_t::relocate_from(): Copies a descriptor that is relocated
to this block, in buf_relocate() and buf_pool_t::shrink(), but keeps
the frame and the modify_clock() of this block. The copy constructor
does not copy the clock: it belongs to the block descriptor and must
only grow during its lifetime, as buf_block_t::modify_clock did.
Otherwise, the clock of this block could return to a value that a
btr_pcur_t saved before the block was freed, and
btr_pcur_t::restore_position() could reuse a record pointer although
the page was modified in another block in the meantime.

buf_page_t::access_time: Store the 16-bit buf_pool.access_clock
rather than a 32-bit millisecond ut_time_ms(). It wraps around every
18.2 hours; ages are computed as uint16_t(now - access_time), which
stays correct across that wrap. This avoids any alignment loss: the
adjacent fields modify_clock_low, modify_clock_high, access_time of
32+16+16 bits nicely add up to 64 bits. access_time is stamped on the
first access after the block was initialized, and on each promotion
by make_young_if_needed(), so that the age of a frequently promoted
block stays exact across the wrap.

buf_pool_t::access_clock: uint16_t(my_interval_timer() / 1000000000),
never 0 (see buf_pool_t::now()), refreshed about once per second by
buf_pool_t::refresh_clock() so that page accesses need not read the
system clock. srv_master_callback() refreshes it. From
buf_pool_t::create() until srv_master_timer is started, and for good
when srv_master_timer is not started (innodb_read_only,
innodb_force_recovery>=2, mariadb-backup), a separate
buf_pool_clock_timer refreshes it.

buf_pool_t::access_clock, buf_pool_t::LRU_old_threshold: Located in a
cache line of their own, because they are read on page accesses and
the adjacent buf_pool fields are frequently written.

buf_page_t::touch(): Stamp access_time on the first access, then
invoke touch_no_stamp(). Return whether this was not the first access,
as the result of buf_page_make_young_if_needed() used to be. A
buffer-fix is sufficient, as in MVCC undo page lookups.

buf_page_t::touch_no_stamp(): Set the PROMOTE flag if
innodb_old_blocks_time is 0, or if accessed_at() is at least that old,
at any position of the block in buf_pool.LRU. Because
buf_pool.access_clock advances once per refresh_clock(), about once per
second, a difference of threshold ticks can span less than
innodb_old_blocks_time; the difference must exceed the threshold, so
that a burst of accesses that crosses a refresh is not promoted. Like
btr_cur_nonleaf_make_young(), do not stamp access_time. Once PROMOTE is
set, later accesses only load the state and return.

buf_page_t::make_young_if_needed(): If PROMOTE is set, invoke
buf_page_t::make_young(). This part is inline, so that a sweep pays no
function call for a block that stays in place. It reads the state of
the block once: buf_pool.mutex, which the caller holds, protects OLD,
and only make_young() clears PROMOTE. The block need not be in
buf_pool.LRU_old: while buf_pool.LRU is shorter than
BUF_LRU_OLD_MIN_LEN, buf_pool.LRU_old does not exist, and blocks are
evicted without ever becoming old. The template parameter only_old
leaves a block that is not old in place; buf_do_flush_list_batch()
uses it, because it visits blocks in the order of oldest_modification,
and moving a block that is not old would not save it from eviction.
The template parameter count_not_young selects whether an old block
that was accessed but not flagged is counted in
buf_pool.stat.n_pages_not_made_young; only the eviction sweeps
buf_LRU_free_from_common_LRU_list() and buf_flush_LRU_list_batch() do
this, and a block that a sweep leaves in the list can be counted again
by a later sweep. As before, buf_pool.stat.n_pages_made_young counts
only the moves of old blocks. Because a sweep moves the block, not the
access, Innodb_buffer_pool_pages_made_young can differ from before for
the same workload.

buf_page_t::make_young(): Clear PROMOTE, stamp access_time, and move
the block to the "recently used" end of buf_pool.LRU. The block can be
read-fixed, because buf_pool_t::unzip() copies PROMOTE and OLD from the
compressed-only descriptor and releases buf_pool.mutex during
buf_zip_decompress(). Unlike buf_page_make_young(), we do not skip such
a block: a page access no longer acquires buf_pool.mutex, and the
sweeps hold buf_pool.mutex but no buffer-fix on the block.

buf_LRU_remove_block(): When buf_pool.LRU becomes too short for
buf_pool.LRU_old, clear the OLD flags from buf_pool.LRU_old to the end
of the list, which holds all the old blocks, not the flags of every
block. While buf_pool.LRU_old does not exist, no block in buf_pool.LRU
is old, so on later removals the loop does nothing.

buf_LRU_add_block(): Write the OLD flag only if it changes. In a list
that is too short for buf_pool.LRU_old, assert that buf_pool.LRU_old
is not defined and that the block is not old, instead of writing the
flag.

innodb_old_blocks_time_update(): New sysvar update callback,
replacing a NULL one, that calls buf_pool_t::set_old_threshold_ms(),
so that SET GLOBAL innodb_old_blocks_time also updates the
LRU_old_threshold in seconds that page accesses read. The threshold
is clamped to 65534 seconds (UINT16_MAX - 1), because the age
uint16_t(now - access_time) cannot exceed 65535 and a larger threshold
would prevent any promotion. The sysvar itself still accepts up to
UINT_MAX32 milliseconds.

buf_pool_t::set_old_threshold_ms(): Round the millisecond threshold
up, not down, to the nearest second. innodb_old_blocks_time is
documented and accepted in milliseconds; flooring instead of ceiling
would make any configured value from 1 to 999 silently behave as 0
(disabled).

buf_page_t::is_accessed(): Renamed accessed_at(), to stop reading as
a boolean. It returns the access_time stamp of the first access or
of the last promotion, in seconds;
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).ACCESS_TIME now reflects
that.

buf_read_ahead_random(): A page now qualifies once accessed_at()
holds, together with either zip.is_promote() or !zip.old().

buf_read_ahead_linear(): Compare access_time stamps as a signed
16-bit difference, not raw unsigned, so the monotonic-access check
stays correct across the access_time wrap. The resolution of the
stamps is 1 second instead of 1 millisecond.

buf_flush_page_cleaner(): Refresh abstime before proceeding to LRU
eviction after an idle period, so that the next my_cond_timedwait()
will not return immediately on a stale deadline.

buf_LRU_scan_and_free_block(): Declare static.

buf_flush_LRU_list_batch(): In a run of blocks that
make_young_if_needed() moves, release and reacquire buf_pool.mutex
after every 512 scanned blocks, except on the first scanned block.
A move costs much less than an eviction, so the stride is longer than
the one of the eviction path.

buf_pool_invalidate(): Define in the same compilation unit with
buf_LRU_scan_and_free_block().

buf_pool.LRU_old_time_threshold: Replaces buf_LRU_old_threshold_ms.

PageConverter::run(): Renamed from fil_iterate(). In debug builds,
initialize and acquire a dummy exclusive latch on the block, so the
assertion in buf_page_t::invalidate() is satisfied; free the latch on
every path, not only on success.

AbstractCallback::m_zip_ssize: Replaces m_zip_size.

page_zip_des_t: Add calc_ssize()/zip_size() helpers, replacing the
zip_size<->ssize conversion duplicated across buf0buf.cc, buf0rea.cc
and row0import.cc.

innodb.buf_lru_scan_resistance: A new big test that checks that pages
of a table that is accessed in one burst are evicted, even if a sweep
reaches them only after innodb_old_blocks_time, that pages accessed
again after that time are promoted, and that pages read while the
buffer pool is being filled are not promoted later.

innodb_zip.n_blobs_700: A new test that stores 700 BLOB pointers on
one ROW_FORMAT=COMPRESSED page, within the 10-bit n_blobs field. The
index_page_splits counter confirms that the clustered index consists
of its root page alone, without depending on which pages remain in the
buffer pool.

Co-Authored-By: Alessandro Vetere <[email protected]>
Mohammad Tafzeel Shams
MDEV-41242 : Fix resource leaks on InnoDB/mariabackup error paths found by Infer

Several error-handling paths returned without releasing a resource
already acquired earlier in the function, or checked the wrong handle
entirely, risking use of an unopened handle.

Changes:
- SysTablespace::read_lsn_and_check_flags(): close the datafile handle
  on header-validation failure.
- xb_process_datadir(): check the freshly opened `dir` handle instead
  of the stale `dbdir`, fixing a handle leak and a possible use of an
  unopened directory handle.
- wsrep.cc / xb_load_list_file(): close file handles before die(), and
  null-check fopen() results in wsrep.cc.
- datadir_iter_new(): free datadir_path and destroy the mutex on the
  os_file_opendir() failure path.
Aleksey Midenkov
MDEV-41157 CREATE DATABASE COMMENT overflows db.opt comment buffer

Bug 1: put_dbopt() used strmov() to copy schema_comment into a fixed
DATABASE_COMMENT_MAXLEN+1 buffer. validate_comment_length() only
truncates comment->length in non-strict sql_mode, leaving comment->str
NUL-terminated at its original (unbounded) length. strmov() copies
until the source NUL, ignoring the truncated length, overflowing the
destination buffer for long comments.

The fix uses strmake() bounded by comment->length instead, matching
the LEX_CSTRING contract (length is authoritative, str need not be
NUL-terminated at length).

Bug 2: write_db_opt() used strxnmov() to copy the full un-truncated
comment until it ran out of buffer space mid-string with no trailing
newline. Which made load_db_opt() silently discard the whole
unterminated "comment=" line on the next restart, losing the comment
entirely instead of just truncating it.

The fix bounds the comment copy into db.opt by the already-validated
comment->length via strmake(), instead of relying on the source
string's own NUL terminator, matching the put_dbopt() fix.

Bug 3: validate_comment_length() only runs on a COMMENT clause given
in the current statement. ALTER DATABASE without one instead pulls
the existing comment off disk via load_db_opt(), which never bounded
it. That unvalidated length then reached write_db_opt()'s own
comment= copy into its stack buffer, so a legacy or hand-edited
db.opt with an overlong comment= line overflowed it on ALTER DATABASE.

The fix: load_db_opt() now clamps the parsed comment to
DATABASE_COMMENT_MAXLEN right when it reads the "comment=" line, so
every consumer (put_dbopt(), write_db_opt()'s ALTER path) always sees
an already-bounded value.

The clamp itself must truncate by bytes, not characters:
Well_formed_prefix()'s LEX_CSTRING overload takes a character count,
but DATABASE_COMMENT_MAXLEN sizes the buffers in bytes.

Bug 4: write_db_opt() still trusted validate_comment_length() to bound
a directly-given COMMENT to DATABASE_COMMENT_MAXLEN bytes, but it only
bounds it to DATABASE_COMMENT_MAXLEN *characters* -- so a multi-byte
comment could still overflow the same fixed buffers.

The fix: write_db_opt() clamps schema_comment to
DATABASE_COMMENT_MAXLEN bytes itself, right after
validate_comment_length() returns.
Raghunandan Bhat
MDEV-31284: SIGSEGV in VDec2_lazy::VDec2_lazy | Item_func_plus::decimal_op

Problem:
  Deeply nested arithmetic expressions can cause a stack overflow
  during execution because `fix_fields` doesn't account for the heavier
  stack frames of the evaluation path (especially decimal operations).

Fix:
  Add `Item_func::m_exec_stack_size` to track type-specific evaluation
  costs. Perform a second check for stack overrun in `fix_fields`
  simulate execution stack depth.
Jan Lindström
MDEV-41075 : Assertion .thd->in_active_multi_stmt in TOI

Clear OPTION_NOT_AUTOCOMMIT/OPTION_BEGIN before opening the table below,
not after. A storage engine may register itself into the "all"
transaction as part of the table open/lock (e.g. InnoDB's
external_lock()), depending on those bits. Clearing them only after the
table is open is too late: the engine has already registered into "all"
using the still-set bits, and since record_gtid is meant to be a standalone
autocommit-style write, nothing will later issue the matching "all"-level
commit to clear that registration and the performance-schema transaction
handle, and it leaks into whatever runs on this THD next.
Aleksey Midenkov
MDEV-29483 Heap-use-after-free (Binary_string::copy()) with window functions

JOIN::make_aggr_tables_info(): a query with a window function buffers
join rows into a postjoin-aggr temporary table via copy_fields().
BLOB/TEXT fields were copied by Copy_field::do_field_eq(), a raw
memcpy of the packed record representation (length + pointer),
leaving the tmp table's field pointing at whatever storage backed
the source row. Once the underlying handler (e.g. InnoDB) frees or
reuses that storage on a later row fetch, any later read of the
buffered blob (e.g. Item_field::str_result() via result_field,
referenced from a WHERE/HAVING subquery predicate) dereferences a
dangling pointer.

The fix forces BLOB/TEXT fields to be deep-copied (do_save_blob())
for any query with window functions, the same way GROUP BY already
does via save_sum_fields.

save_sum_fields controls an unrelated decision (whether to defer
Item_sum computation instead of wiring an incremental result_field,
see the SUM_FUNC_ITEM handling in Create_tmp_table::add_fields())
and must not be conflated with blob-copy safety. Introduce a
separate save_blobs flag, threaded through
JOIN::create_postjoin_aggr_table(), create_tmp_table() and
Create_tmp_table, that independently controls the Copy_field::set()
deep-copy decision.

Create_tmp_table: add a save_blobs constructor parameter, plus an
overload preserving the historical combined behavior (save_blobs
same as save_sum_fields) for existing callers.
Khaled Riyad
MDEV-40551 Copy/Paste friendly output format for MariaDB Command Line Client

Copy/paste friendly output was only reachable by starting the client with
--silent --skip-column-names, which cannot be done from a running
interactive session.

Add \S, a statement terminator which prints the result of one statement in
the tab separated format without column names.

com_silent() sets output_plain, opt_silent and column_names around
com_go(), then restores them, the same way com_ego() handles vertical.
output_plain selects print_tab_data() ahead of the vertical and table
branches, so \S gives the same output whether the session was started
plainly or with --table, --vertical or --silent. --html and --xml still
win, matching \G.
Marko Mäkelä
MDEV-40950: Crash after memory pressure event

buf_pool_t::garbage_collect(): Correctly handle the SHRINK_ABORT
return value of buf_pool_t::shrink().
Marko Mäkelä
fixup! 33ec54440d8b0ebd6fb4a2286c1167b851b1cdce
Marko Mäkelä
MDEV-41359 fix
Georgi (Joro) Kodinov
Added a Contribution Interaction Etiquette section.
Vladislav Vaintroub
MDEV-41008: Fix X509 issuer/subject comparison for OpenSSL 3

OpenSSL 3 escapes '/' and '+' in X509_NAME_oneline() output; OpenSSL 1.1
and WolfSSL don't. A REQUIRE ISSUER/SUBJECT grant from one library can
stop matching after switching to another.

Default comparison stays strcmp(). old_mode=X509_LENIENT_COMPARE opts
into falling back to an escape-aware compare, applied to whichever side
the currently-linked library's own escaping affects, at the cost of
reopening the ambiguity a crafted certificate could exploit to
impersonate another identity.

Adds regression tests against a real certificate with an ambiguous CN

Assisted-by: Claude Sonnet 5 <[email protected]>
Dave Gosselin
MDEV-41212:  multi_source.status_vars fails on MacOS platform

Replace the two recorded reads of Slave_received_heartbeats with an
assertion that the counter is nonzero.

The counter advances once per heartbeat period for as long as the
connection is running.  The test waited for it to reach 2 and then
read it again in a separate query, so a heartbeat arriving between
those two queries recorded an unexpected value.

The same wait timed out when the counter was already past 2 at the
first poll, so it now accepts any value at or above the target.
Thirunarayanan Balathandayuthapani
MDEV-41055 innodb_encryption_threads=0 hangs indefinitely when rotation IOPS is zero

Problem:
=======
  When innodb_encryption_rotation_iops=0, an encryption thread
could be waiting on fil_crypt_iops_cond in fil_crypt_alloc_iops().
fil_crypt_set_thread_cnt() lowers srv_n_fil_crypt_threads and
broadcasts only fil_crypt_thread_cond, so that the waiting thread
never re-evaluates should_shutdown() and never exits.

Solution:
=========
fil_crypt_set_thread_cnt(): Broadcast fil_crypt_iops_cond as well,
so a thread waiting for IOPS wakes up and sees should_shutdown(),
exits.
Thirunarayanan Balathandayuthapani
MDEV-40324 use-of-uninitialized-value after creation of FULLTEXT table failure

Problem:
=======
For fulltext index, row_create_index_for_mysql() calls
fts_create_index_tables(). If creating FTS auxiliary table fails,
error handling performs trx->rollback() of the dictionary
transaction. Rollback removes the parent table from
dictionary cache and frees it. After that,
convert_error_code_to_mysql() reads table->flags after table->heap.
This leads to read of freed memory.

Solution:
========
create_index(): Read table->flags into a local variable before
calling row_create_index_for_mysql()
Marko Mäkelä
MDEV-40950: Crash after memory pressure event

buf_pool_t::garbage_collect(): Correctly handle the SHRINK_ABORT
return value of buf_pool_t::shrink().
Marko Mäkelä
MDEV-33966: buf_page_make_young() is a contention point

The buf_pool.LRU list needs to reasonably accurately reflect recently
accessed blocks, so that they will not be evicted prematurely.
Because the list is protected by buf_pool.mutex, it is not a good idea
to maintain the position on every page access.

Instead of maintaining the LRU position on each access, we will
decide on each access whether the block qualifies for promotion, and
record the decision in a "promote" flag in the block descriptor. The
buf_flush_page_cleaner() thread as well as some traversal of the
buf_pool.LRU list, all of which hold buf_pool.mutex anyway, will move
the flagged blocks to the "recently used" end of buf_pool.LRU.

A block qualifies only if it is accessed again at least
innodb_old_blocks_time after the reference point, which is its first
access and, after that, its most recent promotion. The age is measured
in whole seconds and checked on the access, not when a sweep reaches
the block, so that a block that a table scan accessed in one burst
will not be promoted, however long it stays in buf_pool.LRU_old. After
a promotion, the age is measured from the time the sweep moved the
block, which can be later than the access that qualified it. The rule
applies at any position in buf_pool.LRU, and the flag stays set until
a sweep moves the block. Thus, a block that qualified outside
buf_pool.LRU_old keeps the flag until a sweep reaches it, which is
usually after it has moved into buf_pool.LRU_old; before, such a
block was moved on any access when freed_page_clock showed that it was
no longer close to the "recently used" end.

buf_page_make_young_if_needed(), buf_page_make_young(),
buf_page_peek_if_too_old(), buf_page_peek_if_young(),
btr_cur_nonleaf_make_young(), buf_page_t::set_accessed(): Replaced by
buf_page_t::touch(), buf_page_t::touch_no_stamp() and
buf_page_t::make_young_if_needed().

buf_pool_t::freed_page_clock, buf_page_t::freed_page_clock: Remove.
This is no longer meaningful in the revised design.
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).FREE_PAGE_CLOCK now
always reports 0. Before the first eviction, buf_LRU_stat_update() now
records statistics intervals, and buf_LRU_evict_from_unzip_LRU() uses
its formula instead of assuming a disk-bound workload.

page_zip_des_t::state: An atomic 16-bit field that will include
the PROMOTE and OLD flags that would more logically belong
to buf_page_t. We maintain them here (along with some
ROW_FORMAT=COMPRESSED specific state that is protected by page
latches) in order to avoid race conditions and unnecessary
memory overhead. m_end moves out to its own field, shrinking the
packed word to 16 bits; n_blobs narrows from 12 bits to 10 to make
room for PROMOTE and OLD, still comfortably above the 744-column
maximum.

The helpers that set or clear one bit of page_zip_des_t::state always
perform the atomic read-modify-write, which is a full barrier on x86.
A caller that can find the bit already in the target state checks it
first: touch_no_stamp() for PROMOTE, buf_LRU_add_block() for OLD, and
page_zip_write_rec() for NONEMPTY. page_zip_des_t::copy_from() and
set_n_blobs_and_empty() skip the fetch_add() when the delta is 0.
page_zip_decompress_clust() counts the BLOB pointers locally and
invokes add_n_blobs() once.

page_zip_des_t::copy_from(): Replaces the copy constructor that
page_zip_copy_recs() invoked. It copies m_end, NONEMPTY and n_blobs
from the source page, but preserves the PROMOTE and OLD flags of the
destination block, because they describe the position of that block
in buf_pool.LRU. Copying OLD from the source would make
buf_pool.LRU_old_len inconsistent with the list. Like
set_n_blobs_and_empty(), it updates the state with a single
fetch_add(), which preserves any concurrent change of PROMOTE or OLD.

buf_page_t::init(): Takes the ROW_FORMAT=COMPRESSED shift size
(ssize) directly and clears the zip descriptor state itself, so
callers (buf_block_t::initialise(), buf_page_init_for_read()) no
longer need a separate page_zip_des_init()/page_zip_set_size() call.

buf_page_t::invalidate(): Replaces buf_block_modify_clock_inc().
Instead of maintaining a 64-bit counter, we will maintain one
comprising 32+16=48 bits, in modify_clock_low,modify_clock_high.
Worst case there will be exactly n<<48 calls to
buf_page_t::invalidate() before some operation such as
btr_pcur_t::restore_position() is executed. Such a count should
be extremely unlikely but not completely impossible. It is worth
noting that the DB_TRX_ID is only 48 bits, and each transaction
start and commit/rollback will consume an identifier.

buf_page_t::modify_clock(): Replaces the read access of
buf_block_t::modify_clock. Assert that the caller is holding
a page latch. Note: because invalidate() and modify_clock() are
protected with buf_pool.mutex or the buf_page_t::lock, there can
be no issue with regard to the atomicity of accessing the 48-bit field.

buf_page_t::relocate_from(): Copies a descriptor that is relocated
to this block, in buf_relocate() and buf_pool_t::shrink(), but keeps
the frame and the modify_clock() of this block. The copy constructor
does not copy the clock: it belongs to the block descriptor and must
only grow during its lifetime, as buf_block_t::modify_clock did.
Otherwise, the clock of this block could return to a value that a
btr_pcur_t saved before the block was freed, and
btr_pcur_t::restore_position() could reuse a record pointer although
the page was modified in another block in the meantime.

buf_page_t::access_time: Store the 16-bit buf_pool.access_clock
rather than a 32-bit millisecond ut_time_ms(). It wraps around every
18.2 hours; ages are computed as uint16_t(now - access_time), which
stays correct across that wrap. This avoids any alignment loss: the
adjacent fields modify_clock_low, modify_clock_high, access_time of
32+16+16 bits nicely add up to 64 bits. access_time is stamped on the
first access after the block was initialized, and on each promotion
by make_young_if_needed(), so that the age of a frequently promoted
block stays exact across the wrap.

buf_pool_t::access_clock: uint16_t(my_interval_timer() / 1000000000),
never 0 (see buf_pool_t::now()), refreshed about once per second by
buf_pool_t::refresh_clock() so that page accesses need not read the
system clock. srv_master_callback() refreshes it. From
buf_pool_t::create() until srv_master_timer is started, and for good
when srv_master_timer is not started (innodb_read_only,
innodb_force_recovery>=2, mariadb-backup), a separate
buf_pool_clock_timer refreshes it.

buf_pool_t::access_clock, buf_pool_t::LRU_old_threshold: Located in a
cache line of their own, because they are read on page accesses and
the adjacent buf_pool fields are frequently written.

buf_page_t::touch(): Stamp access_time on the first access, then
invoke touch_no_stamp(). Return whether this was not the first access,
as the result of buf_page_make_young_if_needed() used to be. A
buffer-fix is sufficient, as in MVCC undo page lookups.

buf_page_t::touch_no_stamp(): Set the PROMOTE flag if
innodb_old_blocks_time is 0, or if accessed_at() is at least that old,
at any position of the block in buf_pool.LRU. Because
buf_pool.access_clock advances once per refresh_clock(), about once per
second, a difference of threshold ticks can span less than
innodb_old_blocks_time; the difference must exceed the threshold, so
that a burst of accesses that crosses a refresh is not promoted. Like
btr_cur_nonleaf_make_young(), do not stamp access_time. Once PROMOTE is
set, later accesses only load the state and return.

buf_page_t::make_young_if_needed(): If PROMOTE is set, invoke
buf_page_t::make_young(). This part is inline, so that a sweep pays no
function call for a block that stays in place. It reads the state of
the block once: buf_pool.mutex, which the caller holds, protects OLD,
and only make_young() clears PROMOTE. The block need not be in
buf_pool.LRU_old: while buf_pool.LRU is shorter than
BUF_LRU_OLD_MIN_LEN, buf_pool.LRU_old does not exist, and blocks are
evicted without ever becoming old. The template parameter only_old
leaves a block that is not old in place; buf_do_flush_list_batch()
uses it, because it visits blocks in the order of oldest_modification,
and moving a block that is not old would not save it from eviction.
The template parameter count_not_young selects whether an old block
that has an accessed_at() stamp but is not flagged is counted in
buf_pool.stat.n_pages_not_made_young. This includes a block that had
no access since its last promotion, so the count is of sweep visits,
not of rejected accesses; only the eviction sweeps
buf_LRU_free_from_common_LRU_list() and buf_flush_LRU_list_batch() do
this, and a block that a sweep leaves in the list can be counted again
by a later sweep. As before, buf_pool.stat.n_pages_made_young counts
only the moves of old blocks. Because a sweep moves the block, not the
access, and counts the blocks that it does not move,
Innodb_buffer_pool_pages_made_young and
Innodb_buffer_pool_pages_made_not_young can differ from before for the
same workload.

buf_page_t::make_young(): Clear PROMOTE, stamp access_time, and move
the block to the "recently used" end of buf_pool.LRU. The block can be
read-fixed, because buf_pool_t::unzip() copies PROMOTE and OLD from the
compressed-only descriptor and releases buf_pool.mutex during
buf_zip_decompress(). Unlike buf_page_make_young(), we do not skip such
a block: a page access no longer acquires buf_pool.mutex, and the
sweeps hold buf_pool.mutex but no buffer-fix on the block.

buf_LRU_remove_block(): When buf_pool.LRU becomes too short for
buf_pool.LRU_old, clear the OLD flags from buf_pool.LRU_old to the end
of the list, which holds all the old blocks, not the flags of every
block. While buf_pool.LRU_old does not exist, no block in buf_pool.LRU
is old, so on later removals the loop does nothing.

buf_LRU_add_block(): Write the OLD flag only if it changes. In a list
that is too short for buf_pool.LRU_old, assert that buf_pool.LRU_old
is not defined and that the block is not old, instead of writing the
flag.

innodb_old_blocks_time_update(): New sysvar update callback,
replacing a NULL one, that calls buf_pool_t::set_old_threshold_ms(),
so that SET GLOBAL innodb_old_blocks_time also updates the
LRU_old_threshold in seconds that page accesses read. The threshold
is clamped to 65534 seconds (UINT16_MAX - 1), because the age
uint16_t(now - access_time) cannot exceed 65535 and a larger threshold
would prevent any promotion. The sysvar itself still accepts up to
UINT_MAX32 milliseconds.

buf_pool_t::set_old_threshold_ms(): Round the millisecond threshold
up, not down, to the nearest second. innodb_old_blocks_time is
documented and accepted in milliseconds; flooring instead of ceiling
would make any configured value from 1 to 999 silently behave as 0
(disabled).

buf_page_t::is_accessed(): Renamed accessed_at(), to stop reading as
a boolean. It returns the access_time stamp of the first access or
of the last promotion, in seconds;
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).ACCESS_TIME now reflects
that.

buf_read_ahead_random(): A page now qualifies once accessed_at()
holds, together with either zip.is_promote() or !zip.old().

buf_read_ahead_linear(): Compare access_time stamps as a signed
16-bit difference, not raw unsigned, so the monotonic-access check
stays correct across the access_time wrap. The resolution of the
stamps is 1 second instead of 1 millisecond.

buf_flush_page_cleaner(): Refresh abstime before proceeding to LRU
eviction after an idle period, so that the next my_cond_timedwait()
will not return immediately on a stale deadline.

buf_LRU_scan_and_free_block(): Declare static.

buf_flush_LRU_list_batch(): In a run of blocks that
make_young_if_needed() moves, release and reacquire buf_pool.mutex
after every 512 scanned blocks, except on the first scanned block.
A move costs much less than an eviction, so the stride is longer than
the one of the eviction path.

buf_pool_invalidate(): Define in the same compilation unit with
buf_LRU_scan_and_free_block().

buf_pool.LRU_old_time_threshold: Replaces buf_LRU_old_threshold_ms.

PageConverter::run(): Renamed from fil_iterate(). In debug builds,
initialize and acquire a dummy exclusive latch on the block, so the
assertion in buf_page_t::invalidate() is satisfied; free the latch on
every path, not only on success.

AbstractCallback::m_zip_ssize: Replaces m_zip_size.

page_zip_des_t: Add calc_ssize()/zip_size() helpers, replacing the
zip_size<->ssize conversion duplicated across buf0buf.cc, buf0rea.cc
and row0import.cc.

innodb.buf_lru_scan_resistance: A new big test that checks that pages
of a table that is accessed in one burst are evicted, even if a sweep
reaches them only after innodb_old_blocks_time, that pages accessed
again after that time are promoted, and that pages read while the
buffer pool is being filled are not promoted later.

innodb_zip.n_blobs_700: A new test that stores 700 BLOB pointers on
one ROW_FORMAT=COMPRESSED page, within the 10-bit n_blobs field. The
index_page_splits counter confirms that the clustered index consists
of its root page alone, without depending on which pages remain in the
buffer pool.

Co-Authored-By: Alessandro Vetere <[email protected]>
Rucha Deodhar
MDEV-39049: Memory corruption & crash in check_key_in_list upon using
JSON_KEYS after modifying character set name/collation

Analysis:
Since the length of string is 0, accessing out of boundry memory,
leads to crash

Fix:
If string length is empty, return success.