Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
Oleksandr Byelkin
MDEV-41175 Fix stack-buffer-overflow in backup_log_ddl()

backup_log_ddl() sized its stack log buffer assuming each
identifier is at most ~40 bytes, but add_name_to_buffer() can
expand each identifier character to 5 bytes when re-encoding it
into my_charset_filename, so a RENAME TABLE with long, special-
character names overflowed the buffer (reported under ASAN).

Fixed by sizing the buffer to the true worst case per identifier,
and by making add_str_to_buffer() and its callers assert on an
explicit end-of-buffer pointer before writing.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Kristian Nielsen
Binlog-in-engine: fix couple typos

Signed-off-by: Kristian Nielsen <[email protected]>
Alexey Yurchenko
MDEV-38920 MTR tests for Galera-side fixes

MDEV-38920-evs-config-warn checks that there is a warning about bad
configuration values and they are not accepted.
MDEV-38920-install-timer-expired reproduces 'install timer expired'
situation.
Both tests require fixed Galera library to pass (4.29) hence will be
skipped when run with older versions.
Vladislav Vaintroub
MDEV-40967 PROXY protocol host check sent in clear text mid-SSL handshake

Defer the host-privileged/host-blocked check for a PROXY-header-derived
address until after the client's SSL handshake completes, instead of
sending it immediately in clear text.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Thirunarayanan Balathandayuthapani
MDEV-40319 Instant ALTER TABLE rollback corrupts virtual column

Problem:
=======
ha_innobase_inplace_ctx::~ha_innobase_inplace_ctx() runs, whenever
ctx->instant_table is set and frees the old_v_cols exist also.
old_v_cols and old_n_v_cols are captured in the constructor as
prebuilt_arg->table->v_cols and n_v_cols, i.e. an alias of the live
table's own virtual columns, not a copy. By the time this destructor
runs, old_table->v_cols is either still that same array. If the
failure happened before ctx->instant_column() ever ran like during
prepare_inplace_alter_table_dict() or failure happened during the
commit phase innobase_instant_try(). In both cases old_v_cols is
the table's current, live v_cols array, so this loop destructs
dict_v_col_t objects that are still in use.

Solution:
========
ha_innobase_inplace_ctx::~ha_innobase_inplace_ctx(): Destruct
instant_table->v_cols[], not old_v_cols[]. instant_table is the
independently allocated dict_table_t that prepare_instant() built;
It owns its own v_cols array, whose dict_v_col_t::v_indexes must
be destructed before dict_mem_table_free() reclaims instant_table's memory.
Dave Gosselin
MDEV-35845:  Propagate a constant into an IN predicate

SELECT * FROM t1 WHERE v IN ('a','b') AND v = 'b' kept both conjuncts
when v is a string column, while the equivalent form written with OR
was simplified to v = 'b'.

Two mechanisms can perform a rewrite.  Multiple equalities handle it
when check_simple_equality() builds an Item_equal, which it does only
if the field's charset allows constant propagation.  Up through 10.5 the
default character set was latin1 whose collation handler supports
constant propagation.  MDEV-19123 made utf8mb4 the default in 11.6, and
the utf8 collation handlers report that they do not support constant
propagation.

The other mechanism is propagate_cond_constants(), which rewrote
the OR form under every collation.  It descends through
change_cond_ref_to_const(), which returns on any node whose
eq_cmp_result() is COND_OK.  Item_func_in inherits that value, so the
IN predicate was skipped.

Implement an optimization in change_cond_ref_to_const() that replaces
the predicant of an IN predicate with the constant from an equality at
the same AND level.  The predicant is compared against every value of
the list, so the existing per-operand test from MDEV-7152 is applied
once for each of them.

Only a predicant whose arguments were all aggregated to one comparison
data type is replaced.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Oleksandr Byelkin
MDEV-41258 Crash in Item_func_nextval on prepared re-execution

Sequence prelocking entries were matched across re-executions by
db.str/table_name.str pointer identity, which aliases memory owned by
the TABLE's own mem_root. Once the owning table was closed and reopened
(FLUSH TABLE, table-cache eviction), those pointers went stale and no
longer matched, leaving a dangling duplicate entry that still got
opened and crashed.
Fix: match entries by owning TABLE_LIST and position instead, and
deep-copy db/table_name onto the statement's own arena once, at entry
creation.

The relink was gated the same way as trigger/routine prelocking
discovery (has_prelocking_list), which stays true forever once a
statement's prelocking set includes a trigger, so it silently stopped
running after the first execution for any such statement.
Fix: moved it into its own function, called unconditionally per table,
scoped to DML statements only (not ALTER or INSERT DELAYED).

A stale entry's linked_table pointer could remain set when its owning
table's own open was skipped in a given execution, or when a
mid-statement open_tables() backoff (close_tables_for_reopen()) closed
the owner without a subsequent relink, so a later open_table() backlink
write could land on already-freed memory.
Fix: clear it in TABLE_LIST::reinit_before_use() and in
close_tables_for_reopen().

set_parameters() overwrote lex->default_used every execution from that
execution's own bound parameters only, losing track of a literal
DEFAULT written in the prepared statement text itself.
Fix: also record default_used at prepare time and OR it in.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Kristian Nielsen
Fix hang on master when disabling semi-sync

There is a global variable global_ack_signal_fd used to signal the receiver
thread to wake up when disabling semi-sync. This variable was cleared to -1
in the Ack_listener destructor, which ran at the end of the
Ack_receiver::run() function without any locking. If the thread was delayed
at that point, it could end up overwriting the new value set by a new
receiver thread. This would leave the server in a state with an invalid
global_ack_signal_fd and could cause a subsequent disable of semisync to
fail due to the wakeup not arriving at the receiver thread.

This was seen as a sporadic failure of the test case
rpl.rpl_semi_sync_cond_var_per_thd.

Fix by not modifying the global in constructor/destructor; instead set and
clear the global explicitly, allowing to clear the fd with proper locking
while the mutex is still being held.

Also fix a missing pthread_join(), which would leak thread descriptors and
allow to start a new receiver thread before the old one shut down fully.

(Either of these two changes fix the hang bug).

Signed-off-by: Kristian Nielsen <[email protected]>
Khaled Riyad
MDEV-40551 Copy/Paste friendly output format for MariaDB Command Line Client

Copy/paste friendly output was only reachable by starting the client with
--silent --skip-column-names, which cannot be done from a running
interactive session.

Add \S, a statement terminator which prints the result of one statement in
the tab separated format without column names.

com_silent() sets output_plain, opt_silent and column_names around
com_go(), then restores them, the same way com_ego() handles vertical.
output_plain selects print_tab_data() ahead of the vertical and table
branches, so \S gives the same output whether the session was started
plainly or with --table, --vertical or --silent. --html and --xml still
win, matching \G.
Rucha Deodhar
MDEV-39049: Memory corruption & crash in check_key_in_list upon using
JSON_KEYS after modifying character set name/collation

Analysis:
Buffer overflow crashes and empty key duplication in check_key_in_list.

Fix:
Checking result buffer validity prevents segmentation faults on empty
strings and malformed inputs while preserving correct key matching behavior.
Khaled Riyad
MDEV-40184 Heap-use-after-free in stored procedure when LEFT(CONCAT(...)) result is reused by LOCATE()

LEFT(), RIGHT() and SUBSTR() returned tmp_value pointing into the
buffer passed by the caller, without owning it. When a merged derived
column is used twice, e.g. LOCATE(a, x, a), the second evaluation of
the same item comes from Item_str_func::val_int(), which passes a
local buffer. tmp_value is re-pointed into that buffer, the buffer is
freed on return, and the first result becomes dangling.

Copy the result into tmp_value instead, as MDEV-32758 did for TRIM().
Alexey Yurchenko
MGL-299 Regression in galera_sst_rsync_encrypt_with_key MTR test

Commit b68e29a9c64 explicitly disabled use of SSL encryption in SST
by setting ssl-mode=DISABLED in the top configuration files.
This test is a backward compatibility test so it relies on the
deduction of ssl-mode from the presence of tkey and tcert params
in [sst] section. Unset ssl-mode in config to allow to derive it
from the presence of tkey and tcert.
Aleksey Midenkov
MDEV-40821 SIGSEGV in Window_funcs_sort::setup

Window_funcs_sort failed accessing win_func->window_spec as window
spec was not defined in the query. Usually this is checked by fixing
item win_func, but it was not done by setup_conds.

Actually, earlier check by vers_setup_conds() must fail on DELETE
HISTORY from non-versioned table, but TABLE_LIST for t1 has no
versioning conditions.

The cause was the parser assigning versioning conditions to wrong
table pointed by last_table() which was already switched to another
table from SYSTEM_TIME expression (cs1).

The fix assigns versioning conditions to the correct table stored to
correspondent_table by delete_single_table branch of the parser.

Same fix applied to delete_single_table_for_period (period_conditions),
with tests added for both cases.

Also DBUG_ASSERT in TABLE::vers_switch_partition() checked
delete_history on last_table(), which is also wrong for this case.
Checking the query_tables is correct, as delete_single_table puts
the DELETE target first there, and tables referenced later in the
statement are appended after it.
Oleksandr Byelkin
MDEV-41258 Crash in Item_func_nextval on prepared re-execution

Sequence prelocking entries were matched across re-executions by
db.str/table_name.str pointer identity, aliasing memory on the TABLE's
own mem_root. Once the owning table was closed and reopened, those
pointers went stale, leaving a dangling duplicate entry that still got
opened and crashed.
Fix: match entries by owning TABLE_LIST and position, deep-copied onto
the statement's own arena.

The relink was gated by has_prelocking_list, which stays true forever
once a trigger is involved, so it silently stopped running after the
first execution for such statements.
Fix: moved it into its own function, called unconditionally per table,
keeping the original gating, scoped to DML statements only.

A stale entry's linked_table pointer could survive its owner's open
being skipped, or a mid-statement open_tables() backoff
(close_tables_for_reopen()), leaving open_table()'s backlink write to
land on freed memory.
Fix: clear it in TABLE_LIST::reinit_before_use() and in
close_tables_for_reopen().

set_parameters() overwrote lex->default_used from each execution's own
bound parameters only, losing a literal DEFAULT written in the
statement text itself.
Fix: also record default_used at prepare time and OR it in.

A mid-execute_loop() reprepare() swaps in a freshly re-parsed lex,
losing "EXECUTE ... USING DEFAULT" already folded into default_used by
set_parameters() for this execution.
Fix: save that bit before reprepare() and OR it back in afterwards.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Vladislav Vaintroub
MDEV-41080 startup code on Windows, remove checks for existing service

They were not necessary, just try to run as service, and fallback to
command line.

Add some diagnostics - unexpected errors from StartServiceCtrlDispatcher
and RegisterServiceCtrlHandler are now reported to Windows event log.

Also use authoritative service name, returned as first argument
in svc_main by service control manager.
Raghunandan Bhat
MDEV-41193: ASAN heap-buffer-overflow in ha_connect::CheckCond after select from Connect table

Problem:
  When CONNECT engine pushes a WHERE clause down to an external table,
  it writes the filter into the work area, without checking how much
  space is left. A large string literal in the WHERE clause can overflow
  the work area allocated by the engine. For ex: if connect_work_size is
  set to 4MB and the string literal in the WHERE clause is larger than
  4MB, it can grow past the allocated work area.

Fix:
  Track the space left in the work area and check it before writing. If
  the filter doesn't fit, drop it instead of writing past the buffer.
bsrikanth-mariadb
MDEV-40837 Fix crash recording opt context with NULL character_set_results

Unlike character_set_client or collation_connection,
character_set_results can be set to NULL. When recording context for
a query in opt_context_store_replay.cc,
character_set_results->cs_name was accessed without first checking
whether character_set_results itself was NULL, causing a crash.

Fix Optimizer_context_recorder::dump_sql_script() to check
character_set_results for NULL before accessing cs_name, writing
"SET character_set_results=NULL;" to the replay script in that case
so replay reproduces the recorded session instead of silently
defaulting to the script's SET NAMES charset.

Added a test for the same, in opt_context_store_stats.test, that
also checks the recorded script contains the NULL SET statement.
Aleksey Midenkov
MDEV-41157 Follow-up test fix: strip version-dependent charset result

Follow-up to 6c391b92055.

SHOW CREATE DATABASE always appends "/*!40100 DEFAULT CHARACTER SET
... */", but the default charset differs by version (latin1 on 11.4,
utf8mb4 on 11.8+). Strip it instead of matching a specific value,
since it is not the subject for the fix.

Also switch to evalp: same execution, but logs "$var" instead of the
substituted value, so the replace_regex/disable_query_log hacks go away.
Marko Mäkelä
MDEV-33966: buf_page_make_young() is a contention point

The buf_pool.LRU list needs to reasonably accurately reflect recently
accessed blocks, so that they will not be evicted prematurely.
Because the list is protected by buf_pool.mutex, it is not a good idea
to maintain the position on every page access.

Instead of maintaining the LRU position on each access, we will
decide on each access whether the block qualifies for promotion, and
record the decision in a "promote" flag in the block descriptor. The
buf_flush_page_cleaner() thread as well as some traversal of the
buf_pool.LRU list, all of which hold buf_pool.mutex anyway, will move
the flagged blocks to the "recently used" end of buf_pool.LRU.

A block qualifies only if it is accessed again at least
innodb_old_blocks_time after the reference point, which is its first
access and, after that, its most recent promotion. The age is measured
in whole seconds and checked on the access, not when a sweep reaches
the block, so that a block that a table scan accessed in one burst
will not be promoted, however long it stays in buf_pool.LRU_old. After
a promotion, the age is measured from the time the sweep moved the
block, which can be later than the access that qualified it. The rule
applies at any position in buf_pool.LRU, and the flag stays set until
a sweep moves the block. Thus, a block that qualified outside
buf_pool.LRU_old keeps the flag until a sweep reaches it, which is
usually after it has moved into buf_pool.LRU_old; before, such a
block was moved on any access when freed_page_clock showed that it was
no longer close to the "recently used" end.

buf_page_make_young_if_needed(), buf_page_make_young(),
buf_page_peek_if_too_old(), buf_page_peek_if_young(),
btr_cur_nonleaf_make_young(), buf_page_t::set_accessed(): Replaced by
buf_page_t::touch(), buf_page_t::touch_no_stamp() and
buf_page_t::make_young_if_needed().

buf_pool_t::freed_page_clock, buf_page_t::freed_page_clock: Remove.
This is no longer meaningful in the revised design.
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).FREE_PAGE_CLOCK now
always reports 0. Before the first eviction, buf_LRU_stat_update() now
records statistics intervals, and buf_LRU_evict_from_unzip_LRU() uses
its formula instead of assuming a disk-bound workload.

page_zip_des_t::state: An atomic 16-bit field that will include
the PROMOTE and OLD flags that would more logically belong
to buf_page_t. We maintain them here (along with some
ROW_FORMAT=COMPRESSED specific state that is protected by page
latches) in order to avoid race conditions and unnecessary
memory overhead. m_end moves out to its own field, shrinking the
packed word to 16 bits; n_blobs narrows from 12 bits to 10 to make
room for PROMOTE and OLD, still comfortably above the 744-column
maximum.

The helpers that set or clear one bit of page_zip_des_t::state always
perform the atomic read-modify-write, which is a full barrier on x86.
A caller that can find the bit already in the target state checks it
first: touch_no_stamp() for PROMOTE, buf_LRU_add_block() for OLD, and
page_zip_write_rec() for NONEMPTY. page_zip_des_t::copy_from() and
set_n_blobs_and_empty() skip the fetch_add() when the delta is 0.
page_zip_decompress_clust() counts the BLOB pointers locally and
invokes add_n_blobs() once.

page_zip_des_t::copy_from(): Replaces the copy constructor that
page_zip_copy_recs() invoked. It copies m_end, NONEMPTY and n_blobs
from the source page, but preserves the PROMOTE and OLD flags of the
destination block, because they describe the position of that block
in buf_pool.LRU. Copying OLD from the source would make
buf_pool.LRU_old_len inconsistent with the list. Like
set_n_blobs_and_empty(), it updates the state with a single
fetch_add(), which preserves any concurrent change of PROMOTE or OLD.

buf_page_t::init(): Takes the ROW_FORMAT=COMPRESSED shift size
(ssize) directly and clears the zip descriptor state itself, so
callers (buf_block_t::initialise(), buf_page_init_for_read()) no
longer need a separate page_zip_des_init()/page_zip_set_size() call.

buf_page_t::invalidate(): Replaces buf_block_modify_clock_inc().
Instead of maintaining a 64-bit counter, we will maintain one
comprising 32+16=48 bits, in modify_clock_low,modify_clock_high.
Worst case there will be exactly n<<48 calls to
buf_page_t::invalidate() before some operation such as
btr_pcur_t::restore_position() is executed. Such a count should
be extremely unlikely but not completely impossible. It is worth
noting that the DB_TRX_ID is only 48 bits, and each transaction
start and commit/rollback will consume an identifier.

buf_page_t::modify_clock(): Replaces the read access of
buf_block_t::modify_clock. Assert that the caller is holding
a page latch. Note: because invalidate() and modify_clock() are
protected with buf_pool.mutex or the buf_page_t::lock, there can
be no issue with regard to the atomicity of accessing the 48-bit field.

buf_page_t::relocate_from(): Copies a descriptor that is relocated
to this block, in buf_relocate() and buf_pool_t::shrink(), but keeps
the frame and the modify_clock() of this block. The copy constructor
does not copy the clock: it belongs to the block descriptor and must
only grow during its lifetime, as buf_block_t::modify_clock did.
Otherwise, the clock of this block could return to a value that a
btr_pcur_t saved before the block was freed, and
btr_pcur_t::restore_position() could reuse a record pointer although
the page was modified in another block in the meantime.

buf_page_t::access_time: Store the 16-bit buf_pool.access_clock
rather than a 32-bit millisecond ut_time_ms(). It wraps around every
18.2 hours; ages are computed as uint16_t(now - access_time), which
stays correct across that wrap. This avoids any alignment loss: the
adjacent fields modify_clock_low, modify_clock_high, access_time of
32+16+16 bits nicely add up to 64 bits. access_time is stamped on the
first access after the block was initialized, and on each promotion
by make_young_if_needed(), so that the age of a frequently promoted
block stays exact across the wrap.

buf_pool_t::access_clock: uint16_t(my_interval_timer() / 1000000000),
never 0 (see buf_pool_t::now()), refreshed about once per second by
buf_pool_t::refresh_clock() so that page accesses need not read the
system clock. srv_master_callback() refreshes it. From
buf_pool_t::create() until srv_master_timer is started, and for good
when srv_master_timer is not started (innodb_read_only,
innodb_force_recovery>=2, mariadb-backup), a separate
buf_pool_clock_timer refreshes it.

buf_pool_t::access_clock, buf_pool_t::LRU_old_threshold: Located in a
cache line of their own, because they are read on page accesses and
the adjacent buf_pool fields are frequently written.

buf_page_t::touch(): Stamp access_time on the first access, then
invoke touch_no_stamp(). Return whether this was not the first access,
as the result of buf_page_make_young_if_needed() used to be. A
buffer-fix is sufficient, as in MVCC undo page lookups.

buf_page_t::touch_no_stamp(): Set the PROMOTE flag if
innodb_old_blocks_time is 0, or if accessed_at() is at least that old,
at any position of the block in buf_pool.LRU. Because
buf_pool.access_clock advances once per refresh_clock(), about once per
second, a difference of threshold ticks can span less than
innodb_old_blocks_time; the difference must exceed the threshold, so
that a burst of accesses that crosses a refresh is not promoted. Like
btr_cur_nonleaf_make_young(), do not stamp access_time. Once PROMOTE is
set, later accesses only load the state and return.

buf_page_t::make_young_if_needed(): If PROMOTE is set, invoke
buf_page_t::make_young(). This part is inline, so that a sweep pays no
function call for a block that stays in place. It reads the state of
the block once: buf_pool.mutex, which the caller holds, protects OLD,
and only make_young() clears PROMOTE. The block need not be in
buf_pool.LRU_old: while buf_pool.LRU is shorter than
BUF_LRU_OLD_MIN_LEN, buf_pool.LRU_old does not exist, and blocks are
evicted without ever becoming old. The template parameter only_old
leaves a block that is not old in place; buf_do_flush_list_batch()
uses it, because it visits blocks in the order of oldest_modification,
and moving a block that is not old would not save it from eviction.
The template parameter count_not_young selects whether an old block
that was accessed but not flagged is counted in
buf_pool.stat.n_pages_not_made_young; only the eviction sweeps
buf_LRU_free_from_common_LRU_list() and buf_flush_LRU_list_batch() do
this, and a block that a sweep leaves in the list can be counted again
by a later sweep. As before, buf_pool.stat.n_pages_made_young counts
only the moves of old blocks. Because a sweep moves the block, not the
access, Innodb_buffer_pool_pages_made_young can differ from before for
the same workload.

buf_page_t::make_young(): Clear PROMOTE, stamp access_time, and move
the block to the "recently used" end of buf_pool.LRU. The block can be
read-fixed, because buf_pool_t::unzip() copies PROMOTE and OLD from the
compressed-only descriptor and releases buf_pool.mutex during
buf_zip_decompress(). Unlike buf_page_make_young(), we do not skip such
a block: a page access no longer acquires buf_pool.mutex, and the
sweeps hold buf_pool.mutex but no buffer-fix on the block.

buf_LRU_remove_block(): When buf_pool.LRU becomes too short for
buf_pool.LRU_old, clear the OLD flags from buf_pool.LRU_old to the end
of the list, which holds all the old blocks, not the flags of every
block. While buf_pool.LRU_old does not exist, no block in buf_pool.LRU
is old, so on later removals the loop does nothing.

buf_LRU_add_block(): Write the OLD flag only if it changes. In a list
that is too short for buf_pool.LRU_old, assert that buf_pool.LRU_old
is not defined and that the block is not old, instead of writing the
flag.

innodb_old_blocks_time_update(): New sysvar update callback,
replacing a NULL one, that calls buf_pool_t::set_old_threshold_ms(),
so that SET GLOBAL innodb_old_blocks_time also updates the
LRU_old_threshold in seconds that page accesses read. The threshold
is clamped to 65534 seconds (UINT16_MAX - 1), because the age
uint16_t(now - access_time) cannot exceed 65535 and a larger threshold
would prevent any promotion. The sysvar itself still accepts up to
UINT_MAX32 milliseconds.

buf_pool_t::set_old_threshold_ms(): Round the millisecond threshold
up, not down, to the nearest second. innodb_old_blocks_time is
documented and accepted in milliseconds; flooring instead of ceiling
would make any configured value from 1 to 999 silently behave as 0
(disabled).

buf_page_t::is_accessed(): Renamed accessed_at(), to stop reading as
a boolean. It returns the access_time stamp of the first access or
of the last promotion, in seconds;
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).ACCESS_TIME now reflects
that.

buf_read_ahead_random(): A page now qualifies once accessed_at()
holds, together with either zip.is_promote() or !zip.old().

buf_read_ahead_linear(): Compare access_time stamps as a signed
16-bit difference, not raw unsigned, so the monotonic-access check
stays correct across the access_time wrap. The resolution of the
stamps is 1 second instead of 1 millisecond.

buf_flush_page_cleaner(): Refresh abstime before proceeding to LRU
eviction after an idle period, so that the next my_cond_timedwait()
will not return immediately on a stale deadline.

buf_LRU_scan_and_free_block(): Declare static.

buf_flush_LRU_list_batch(): In a run of blocks that
make_young_if_needed() moves, release and reacquire buf_pool.mutex
after every 512 scanned blocks, except on the first scanned block.
A move costs much less than an eviction, so the stride is longer than
the one of the eviction path.

buf_pool_invalidate(): Define in the same compilation unit with
buf_LRU_scan_and_free_block().

buf_pool.LRU_old_time_threshold: Replaces buf_LRU_old_threshold_ms.

PageConverter::run(): Renamed from fil_iterate(). In debug builds,
initialize and acquire a dummy exclusive latch on the block, so the
assertion in buf_page_t::invalidate() is satisfied; free the latch on
every path, not only on success.

AbstractCallback::m_zip_ssize: Replaces m_zip_size.

page_zip_des_t: Add calc_ssize()/zip_size() helpers, replacing the
zip_size<->ssize conversion duplicated across buf0buf.cc, buf0rea.cc
and row0import.cc.

innodb.buf_lru_scan_resistance: A new big test that checks that pages
of a table that is accessed in one burst are evicted, even if a sweep
reaches them only after innodb_old_blocks_time, that pages accessed
again after that time are promoted, and that pages read while the
buffer pool is being filled are not promoted later.

innodb_zip.n_blobs_700: A new test that stores 700 BLOB pointers on
one ROW_FORMAT=COMPRESSED page, within the 10-bit n_blobs field. The
index_page_splits counter confirms that the clustered index consists
of its root page alone, without depending on which pages remain in the
buffer pool.

Co-Authored-By: Alessandro Vetere <[email protected]>
Mohammad Tafzeel Shams
MDEV-41242 : Fix resource leaks on InnoDB/mariabackup error paths found by Infer

Several error-handling paths returned without releasing a resource
already acquired earlier in the function, or checked the wrong handle
entirely, risking use of an unopened handle.

Changes:
- SysTablespace::read_lsn_and_check_flags(): close the datafile handle
  on header-validation failure.
- xb_process_datadir(): check the freshly opened `dir` handle instead
  of the stale `dbdir`, fixing a handle leak and a possible use of an
  unopened directory handle.
- wsrep.cc / xb_load_list_file(): close file handles before die(), and
  null-check fopen() results in wsrep.cc.
- datadir_iter_new(): free datadir_path and destroy the mutex on the
  os_file_opendir() failure path.
Aleksey Midenkov
MDEV-41157 CREATE DATABASE COMMENT overflows db.opt comment buffer

Bug 1: put_dbopt() used strmov() to copy schema_comment into a fixed
DATABASE_COMMENT_MAXLEN+1 buffer. validate_comment_length() only
truncates comment->length in non-strict sql_mode, leaving comment->str
NUL-terminated at its original (unbounded) length. strmov() copies
until the source NUL, ignoring the truncated length, overflowing the
destination buffer for long comments.

The fix uses strmake() bounded by comment->length instead, matching
the LEX_CSTRING contract (length is authoritative, str need not be
NUL-terminated at length).

Bug 2: write_db_opt() used strxnmov() to copy the full un-truncated
comment until it ran out of buffer space mid-string with no trailing
newline. Which made load_db_opt() silently discard the whole
unterminated "comment=" line on the next restart, losing the comment
entirely instead of just truncating it.

The fix bounds the comment copy into db.opt by the already-validated
comment->length via strmake(), instead of relying on the source
string's own NUL terminator, matching the put_dbopt() fix.

Bug 3: validate_comment_length() only runs on a COMMENT clause given
in the current statement. ALTER DATABASE without one instead pulls
the existing comment off disk via load_db_opt(), which never bounded
it. That unvalidated length then reached write_db_opt()'s own
comment= copy into its stack buffer, so a legacy or hand-edited
db.opt with an overlong comment= line overflowed it on ALTER DATABASE.

The fix: load_db_opt() now clamps the parsed comment to
DATABASE_COMMENT_MAXLEN right when it reads the "comment=" line, so
every consumer (put_dbopt(), write_db_opt()'s ALTER path) always sees
an already-bounded value.

The clamp itself must truncate by bytes, not characters:
Well_formed_prefix()'s LEX_CSTRING overload takes a character count,
but DATABASE_COMMENT_MAXLEN sizes the buffers in bytes.

Bug 4: write_db_opt() still trusted validate_comment_length() to bound
a directly-given COMMENT to DATABASE_COMMENT_MAXLEN bytes, but it only
bounds it to DATABASE_COMMENT_MAXLEN *characters* -- so a multi-byte
comment could still overflow the same fixed buffers.

The fix: write_db_opt() clamps schema_comment to
DATABASE_COMMENT_MAXLEN bytes itself, right after
validate_comment_length() returns.
Raghunandan Bhat
MDEV-31284: SIGSEGV in VDec2_lazy::VDec2_lazy | Item_func_plus::decimal_op

Problem:
  Deeply nested arithmetic expressions can cause a stack overflow
  during execution because `fix_fields` doesn't account for the heavier
  stack frames of the evaluation path (especially decimal operations).

Fix:
  Add `Item_func::m_exec_stack_size` to track type-specific evaluation
  costs. Perform a second check for stack overrun in `fix_fields`
  simulate execution stack depth.
Jan Lindström
MDEV-41075 : Assertion .thd->in_active_multi_stmt in TOI

Clear OPTION_NOT_AUTOCOMMIT/OPTION_BEGIN before opening the table below,
not after. A storage engine may register itself into the "all"
transaction as part of the table open/lock (e.g. InnoDB's
external_lock()), depending on those bits. Clearing them only after the
table is open is too late: the engine has already registered into "all"
using the still-set bits, and since record_gtid is meant to be a standalone
autocommit-style write, nothing will later issue the matching "all"-level
commit to clear that registration and the performance-schema transaction
handle, and it leaks into whatever runs on this THD next.
Aleksey Midenkov
MDEV-29483 Heap-use-after-free (Binary_string::copy()) with window functions

JOIN::make_aggr_tables_info(): a query with a window function buffers
join rows into a postjoin-aggr temporary table via copy_fields().
BLOB/TEXT fields were copied by Copy_field::do_field_eq(), a raw
memcpy of the packed record representation (length + pointer),
leaving the tmp table's field pointing at whatever storage backed
the source row. Once the underlying handler (e.g. InnoDB) frees or
reuses that storage on a later row fetch, any later read of the
buffered blob (e.g. Item_field::str_result() via result_field,
referenced from a WHERE/HAVING subquery predicate) dereferences a
dangling pointer.

The fix forces BLOB/TEXT fields to be deep-copied (do_save_blob())
for any query with window functions, the same way GROUP BY already
does via save_sum_fields.

save_sum_fields controls an unrelated decision (whether to defer
Item_sum computation instead of wiring an incremental result_field,
see the SUM_FUNC_ITEM handling in Create_tmp_table::add_fields())
and must not be conflated with blob-copy safety. Introduce a
separate save_blobs flag, threaded through
JOIN::create_postjoin_aggr_table(), create_tmp_table() and
Create_tmp_table, that independently controls the Copy_field::set()
deep-copy decision.

Create_tmp_table: add a save_blobs constructor parameter, plus an
overload preserving the historical combined behavior (save_blobs
same as save_sum_fields) for existing callers.
Oleksandr Byelkin
MDEV-41144 Assertion `scale <= precision' failed in decimal_bin_size

dynamic_column_decimal_read() derived intg/frac from an unbounded
var-uint and only checked scale>precision in int domain, so a crafted
value could still overflow decimal_bin_size()'s uint16 parameters,
tripping its scale<=precision assert (an out-of-bounds array read in
release builds).

Bound intg/frac against DECIMAL_MAX_POSSIBLE_PRECISION before use, and
reject the var-uint decoder's error signal, as the sibling
dynamic_column_string_read() already does. Validate the decoded
ulonglong before narrowing it to int, so a value wrapping around on
narrowing cannot bypass the bound.
Khaled Riyad
MDEV-40551 Copy/Paste friendly output format for MariaDB Command Line Client

Copy/paste friendly output was only reachable by starting the client with
--silent --skip-column-names, which cannot be done from a running
interactive session.

Add \S, a statement terminator which prints the result of one statement in
the tab separated format without column names.

com_silent() sets output_plain, opt_silent and column_names around
com_go(), then restores them, the same way com_ego() handles vertical.
output_plain selects print_tab_data() ahead of the vertical and table
branches, so \S gives the same output whether the session was started
plainly or with --table, --vertical or --silent. --html and --xml still
win, matching \G.
Marko Mäkelä
MDEV-40950: Crash after memory pressure event

buf_pool_t::garbage_collect(): Correctly handle the SHRINK_ABORT
return value of buf_pool_t::shrink().
bsrikanth-mariadb
MDEV-36986: Support tracing array of primitive types

Json_writer had two separate code paths: add_unquoted_str() for
numbers/bool/null, and add_escaped_str() which added the surrounding
quotes itself for strings. Single_line_formatting_helper, which
buffers consecutive array/object elements to decide if they fit on
one line, assumed only strings could ever be buffered and always
wrapped the flushed values in quotes. As a result, arrays of numbers
(e.g. "depends_on_map_bits", "rec_per_key") were incorrectly rendered
with their elements quoted as strings.

Unify both paths into add_escaped_quoted_str(): the caller now hands
over bytes that are already in their final on-the-wire form. String
escaping (json_escape_to_string) writes its own surrounding quotes,
while numbers/bool/null are passed through unquoted, so the one-line
helper just concatenates the buffered payloads on flush instead of
adding quotes itself. A DBUG_ASSERT in add_escaped_quoted_str() now
checks that every payload is already a quoted string or a bare
number/bool/null token, so a caller that violates the contract trips
an assertion in debug builds instead of silently producing invalid
JSON.

Also:
- Fix Json_writer_array::add(ulonglong)/(size_t), which went through
  add_ll() with a cast to longlong and corrupted large unsigned
  values (e.g. ULLONG_MAX); route them through add_ull() instead.
- Fix mysql-test/include/opt_context_schema.inc: "subquery_runs" was
  nested inside the preceding object instead of being a sibling
  member, and "rec_per_key" items are now declared as "number" to
  match the corrected output.
- Update recorded .result files for opt_trace, opt_context_*, and
  subselect_mat_analyze_json to reflect numbers/booleans no longer
  being quoted inside JSON arrays.
- Extend unittest/sql/my_json_writer-t.cc with coverage for arrays
  of primitives: plain integers, mixed types, sizes, multi-line
  arrays, values flushed before a nested object, strings that were
  already escaped by the one-line buffer, and invalid utf8mb4 input
  through add_str().
- Rewrite the "multi-line array of integers" test to actually
  overflow the one-line buffer (9 seven-digit numbers instead of 7),
  since the previous version fit on one line and never exercised the
  element-per-line flush path it claimed to test.
- Drop the now-redundant single-argument add_escaped_quoted_str()
  overload; all call sites already know their length.
- Update json_escape_to_string()'s doc comments (my_json_writer.h,
  sql_json_lib.h) to state that it quotes its output, not just
  escapes it.
Marko Mäkelä
fixup! 33ec54440d8b0ebd6fb4a2286c1167b851b1cdce
Dave Gosselin
MDEV-35747:  Wrong result from prepared TVC with parameter markers

The setup of column type information in table_value_constr::prepare()
was wrapped in an "if (!holders)" guard so that it runs only once per
statement.  However, the guard was too wide because it bound the
allocation of item holders (which should happen only once) to the
collection of type information (which should happen on each execution).

This leaves the TVC stuck with whatever placeholder type the parameter
had when the holders were first built, which may not match the type of
the next substitution.  A parameter marker has no type of its own until
a value is bound at EXECUTE time.  So both the TVC types and the
corresponding Item_type_holder instance in the SELECT item list must be
computed again on every EXECUTE.

Type holder allocation happens on the first call to the prepare()
function but that doesn't always coincide with a PREPARE.  It does for a
prepared statement whose table value constructor comes from the parser.
For a statement of a stored procedure, and for a table value constructor
that the conversion of an IN predicate into an IN subquery creates,
allocation happens instead on the first execution.  The corresponding
assertion allows the first execution and conventional execution as well
as PREPARE.

This patch separates the work done once per statement from the work done
on every execution as described above.

Whether the SELECT list of Item_type_holder instances has been built is
read from that list rather than from the holder array.  An error raised
while collecting the types leaves the array allocated and the list
empty, and the next call has to build the list.  Nullability starts over
on each collection so that it reflects the values of the current
execution.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Marko Mäkelä
MDEV-41359 fix
Georgi (Joro) Kodinov
Added a Contribution Interaction Etiquette section.
Dave Gosselin
MDEV-41303:  rand() in a semi-join subquery is checked on outer rows

Converting an IN subquery to a semi-join moves its WHERE into the
parent WHERE.  A condition there such as rand(1) < 0.09 reads no table,
so it is attached to the last table of the join order that is outside
any materialized semi-join.  With SJ-Materialization it was checked once
for each outer row instead of once for each row of the subquery, and
the query returned a wrong count.

Do not convert a subquery to a semi-join when it contains a function
with a random result (UNCACHEABLE_RAND).  Derived tables already follow
this rule.  ROWNUM sets the same flag, so this replaces the check for
ROWNUM.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Vladislav Vaintroub
MDEV-41008: Fix X509 issuer/subject comparison for OpenSSL 3

OpenSSL 3 escapes '/' and '+' in X509_NAME_oneline() output; OpenSSL 1.1
and WolfSSL don't. A REQUIRE ISSUER/SUBJECT grant from one library can
stop matching after switching to another.

Default comparison stays strcmp(). old_mode=X509_LENIENT_COMPARE opts
into falling back to an escape-aware compare, applied to whichever side
the currently-linked library's own escaping affects, at the cost of
reopening the ambiguity a crafted certificate could exploit to
impersonate another identity.

Adds regression tests against a real certificate with an ambiguous CN

Assisted-by: Claude Sonnet 5 <[email protected]>
Dave Gosselin
MDEV-41212:  multi_source.status_vars fails on MacOS platform

Replace the two recorded reads of Slave_received_heartbeats with an
assertion that the counter is nonzero.

The counter advances once per heartbeat period for as long as the
connection is running.  The test waited for it to reach 2 and then
read it again in a separate query, so a heartbeat arriving between
those two queries recorded an unexpected value.

The same wait timed out when the counter was already past 2 at the
first poll, so it now accepts any value at or above the target.
Thirunarayanan Balathandayuthapani
MDEV-41055 innodb_encryption_threads=0 hangs indefinitely when rotation IOPS is zero

Problem:
=======
  When innodb_encryption_rotation_iops=0, an encryption thread
could be waiting on fil_crypt_iops_cond in fil_crypt_alloc_iops().
fil_crypt_set_thread_cnt() lowers srv_n_fil_crypt_threads and
broadcasts only fil_crypt_thread_cond, so that the waiting thread
never re-evaluates should_shutdown() and never exits.

Solution:
=========
fil_crypt_set_thread_cnt(): Broadcast fil_crypt_iops_cond as well,
so a thread waiting for IOPS wakes up and sees should_shutdown(),
exits.
Thirunarayanan Balathandayuthapani
MDEV-40324 use-of-uninitialized-value after creation of FULLTEXT table failure

Problem:
=======
For fulltext index, row_create_index_for_mysql() calls
fts_create_index_tables(). If creating FTS auxiliary table fails,
error handling performs trx->rollback() of the dictionary
transaction. Rollback removes the parent table from
dictionary cache and frees it. After that,
convert_error_code_to_mysql() reads table->flags after table->heap.
This leads to read of freed memory.

Solution:
========
create_index(): Read table->flags into a local variable before
calling row_create_index_for_mysql()
Marko Mäkelä
MDEV-40950: Crash after memory pressure event

buf_pool_t::garbage_collect(): Correctly handle the SHRINK_ABORT
return value of buf_pool_t::shrink().
Marko Mäkelä
MDEV-33966: buf_page_make_young() is a contention point

The buf_pool.LRU list needs to reasonably accurately reflect recently
accessed blocks, so that they will not be evicted prematurely.
Because the list is protected by buf_pool.mutex, it is not a good idea
to maintain the position on every page access.

Instead of maintaining the LRU position on each access, we will
decide on each access whether the block qualifies for promotion, and
record the decision in a "promote" flag in the block descriptor. The
buf_flush_page_cleaner() thread as well as some traversal of the
buf_pool.LRU list, all of which hold buf_pool.mutex anyway, will move
the flagged blocks to the "recently used" end of buf_pool.LRU.

A block qualifies only if it is accessed again at least
innodb_old_blocks_time after the reference point, which is its first
access and, after that, its most recent promotion. The age is measured
in whole seconds and checked on the access, not when a sweep reaches
the block, so that a block that a table scan accessed in one burst
will not be promoted, however long it stays in buf_pool.LRU_old. After
a promotion, the age is measured from the time the sweep moved the
block, which can be later than the access that qualified it. The rule
applies at any position in buf_pool.LRU, and the flag stays set until
a sweep moves the block. Thus, a block that qualified outside
buf_pool.LRU_old keeps the flag until a sweep reaches it, which is
usually after it has moved into buf_pool.LRU_old; before, such a
block was moved on any access when freed_page_clock showed that it was
no longer close to the "recently used" end.

buf_page_make_young_if_needed(), buf_page_make_young(),
buf_page_peek_if_too_old(), buf_page_peek_if_young(),
btr_cur_nonleaf_make_young(), buf_page_t::set_accessed(): Replaced by
buf_page_t::touch(), buf_page_t::touch_no_stamp() and
buf_page_t::make_young_if_needed().

buf_pool_t::freed_page_clock, buf_page_t::freed_page_clock: Remove.
This is no longer meaningful in the revised design.
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).FREE_PAGE_CLOCK now
always reports 0. Before the first eviction, buf_LRU_stat_update() now
records statistics intervals, and buf_LRU_evict_from_unzip_LRU() uses
its formula instead of assuming a disk-bound workload.

page_zip_des_t::state: An atomic 16-bit field that will include
the PROMOTE and OLD flags that would more logically belong
to buf_page_t. We maintain them here (along with some
ROW_FORMAT=COMPRESSED specific state that is protected by page
latches) in order to avoid race conditions and unnecessary
memory overhead. m_end moves out to its own field, shrinking the
packed word to 16 bits; n_blobs narrows from 12 bits to 10 to make
room for PROMOTE and OLD, still comfortably above the 744-column
maximum.

The helpers that set or clear one bit of page_zip_des_t::state always
perform the atomic read-modify-write, which is a full barrier on x86.
A caller that can find the bit already in the target state checks it
first: touch_no_stamp() for PROMOTE, buf_LRU_add_block() for OLD, and
page_zip_write_rec() for NONEMPTY. page_zip_des_t::copy_from() and
set_n_blobs_and_empty() skip the fetch_add() when the delta is 0.
page_zip_decompress_clust() counts the BLOB pointers locally and
invokes add_n_blobs() once.

page_zip_des_t::copy_from(): Replaces the copy constructor that
page_zip_copy_recs() invoked. It copies m_end, NONEMPTY and n_blobs
from the source page, but preserves the PROMOTE and OLD flags of the
destination block, because they describe the position of that block
in buf_pool.LRU. Copying OLD from the source would make
buf_pool.LRU_old_len inconsistent with the list. Like
set_n_blobs_and_empty(), it updates the state with a single
fetch_add(), which preserves any concurrent change of PROMOTE or OLD.

buf_page_t::init(): Takes the ROW_FORMAT=COMPRESSED shift size
(ssize) directly and clears the zip descriptor state itself, so
callers (buf_block_t::initialise(), buf_page_init_for_read()) no
longer need a separate page_zip_des_init()/page_zip_set_size() call.

buf_page_t::invalidate(): Replaces buf_block_modify_clock_inc().
Instead of maintaining a 64-bit counter, we will maintain one
comprising 32+16=48 bits, in modify_clock_low,modify_clock_high.
Worst case there will be exactly n<<48 calls to
buf_page_t::invalidate() before some operation such as
btr_pcur_t::restore_position() is executed. Such a count should
be extremely unlikely but not completely impossible. It is worth
noting that the DB_TRX_ID is only 48 bits, and each transaction
start and commit/rollback will consume an identifier.

buf_page_t::modify_clock(): Replaces the read access of
buf_block_t::modify_clock. Assert that the caller is holding
a page latch. Note: because invalidate() and modify_clock() are
protected with buf_pool.mutex or the buf_page_t::lock, there can
be no issue with regard to the atomicity of accessing the 48-bit field.

buf_page_t::relocate_from(): Copies a descriptor that is relocated
to this block, in buf_relocate() and buf_pool_t::shrink(), but keeps
the frame and the modify_clock() of this block. The copy constructor
does not copy the clock: it belongs to the block descriptor and must
only grow during its lifetime, as buf_block_t::modify_clock did.
Otherwise, the clock of this block could return to a value that a
btr_pcur_t saved before the block was freed, and
btr_pcur_t::restore_position() could reuse a record pointer although
the page was modified in another block in the meantime.

buf_page_t::access_time: Store the 16-bit buf_pool.access_clock
rather than a 32-bit millisecond ut_time_ms(). It wraps around every
18.2 hours; ages are computed as uint16_t(now - access_time), which
stays correct across that wrap. This avoids any alignment loss: the
adjacent fields modify_clock_low, modify_clock_high, access_time of
32+16+16 bits nicely add up to 64 bits. access_time is stamped on the
first access after the block was initialized, and on each promotion
by make_young_if_needed(), so that the age of a frequently promoted
block stays exact across the wrap.

buf_pool_t::access_clock: uint16_t(my_interval_timer() / 1000000000),
never 0 (see buf_pool_t::now()), refreshed about once per second by
buf_pool_t::refresh_clock() so that page accesses need not read the
system clock. srv_master_callback() refreshes it. From
buf_pool_t::create() until srv_master_timer is started, and for good
when srv_master_timer is not started (innodb_read_only,
innodb_force_recovery>=2, mariadb-backup), a separate
buf_pool_clock_timer refreshes it.

buf_pool_t::access_clock, buf_pool_t::LRU_old_threshold: Located in a
cache line of their own, because they are read on page accesses and
the adjacent buf_pool fields are frequently written.

buf_page_t::touch(): Stamp access_time on the first access, then
invoke touch_no_stamp(). Return whether this was not the first access,
as the result of buf_page_make_young_if_needed() used to be. A
buffer-fix is sufficient, as in MVCC undo page lookups.

buf_page_t::touch_no_stamp(): Set the PROMOTE flag if
innodb_old_blocks_time is 0, or if accessed_at() is at least that old,
at any position of the block in buf_pool.LRU. Because
buf_pool.access_clock advances once per refresh_clock(), about once per
second, a difference of threshold ticks can span less than
innodb_old_blocks_time; the difference must exceed the threshold, so
that a burst of accesses that crosses a refresh is not promoted. Like
btr_cur_nonleaf_make_young(), do not stamp access_time. Once PROMOTE is
set, later accesses only load the state and return.

buf_page_t::make_young_if_needed(): If PROMOTE is set, invoke
buf_page_t::make_young(). This part is inline, so that a sweep pays no
function call for a block that stays in place. It reads the state of
the block once: buf_pool.mutex, which the caller holds, protects OLD,
and only make_young() clears PROMOTE. The block need not be in
buf_pool.LRU_old: while buf_pool.LRU is shorter than
BUF_LRU_OLD_MIN_LEN, buf_pool.LRU_old does not exist, and blocks are
evicted without ever becoming old. The template parameter only_old
leaves a block that is not old in place; buf_do_flush_list_batch()
uses it, because it visits blocks in the order of oldest_modification,
and moving a block that is not old would not save it from eviction.
The template parameter count_not_young selects whether an old block
that has an accessed_at() stamp but is not flagged is counted in
buf_pool.stat.n_pages_not_made_young. This includes a block that had
no access since its last promotion, so the count is of sweep visits,
not of rejected accesses; only the eviction sweeps
buf_LRU_free_from_common_LRU_list() and buf_flush_LRU_list_batch() do
this, and a block that a sweep leaves in the list can be counted again
by a later sweep. As before, buf_pool.stat.n_pages_made_young counts
only the moves of old blocks. Because a sweep moves the block, not the
access, and counts the blocks that it does not move,
Innodb_buffer_pool_pages_made_young and
Innodb_buffer_pool_pages_made_not_young can differ from before for the
same workload.

buf_page_t::make_young(): Clear PROMOTE, stamp access_time, and move
the block to the "recently used" end of buf_pool.LRU. The block can be
read-fixed, because buf_pool_t::unzip() copies PROMOTE and OLD from the
compressed-only descriptor and releases buf_pool.mutex during
buf_zip_decompress(). Unlike buf_page_make_young(), we do not skip such
a block: a page access no longer acquires buf_pool.mutex, and the
sweeps hold buf_pool.mutex but no buffer-fix on the block.

buf_LRU_remove_block(): When buf_pool.LRU becomes too short for
buf_pool.LRU_old, clear the OLD flags from buf_pool.LRU_old to the end
of the list, which holds all the old blocks, not the flags of every
block. While buf_pool.LRU_old does not exist, no block in buf_pool.LRU
is old, so on later removals the loop does nothing.

buf_LRU_add_block(): Write the OLD flag only if it changes. In a list
that is too short for buf_pool.LRU_old, assert that buf_pool.LRU_old
is not defined and that the block is not old, instead of writing the
flag.

innodb_old_blocks_time_update(): New sysvar update callback,
replacing a NULL one, that calls buf_pool_t::set_old_threshold_ms(),
so that SET GLOBAL innodb_old_blocks_time also updates the
LRU_old_threshold in seconds that page accesses read. The threshold
is clamped to 65534 seconds (UINT16_MAX - 1), because the age
uint16_t(now - access_time) cannot exceed 65535 and a larger threshold
would prevent any promotion. The sysvar itself still accepts up to
UINT_MAX32 milliseconds.

buf_pool_t::set_old_threshold_ms(): Round the millisecond threshold
up, not down, to the nearest second. innodb_old_blocks_time is
documented and accepted in milliseconds; flooring instead of ceiling
would make any configured value from 1 to 999 silently behave as 0
(disabled).

buf_page_t::is_accessed(): Renamed accessed_at(), to stop reading as
a boolean. It returns the access_time stamp of the first access or
of the last promotion, in seconds;
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).ACCESS_TIME now reflects
that.

buf_read_ahead_random(): A page now qualifies once accessed_at()
holds, together with either zip.is_promote() or !zip.old().

buf_read_ahead_linear(): Compare access_time stamps as a signed
16-bit difference, not raw unsigned, so the monotonic-access check
stays correct across the access_time wrap. The resolution of the
stamps is 1 second instead of 1 millisecond.

buf_flush_page_cleaner(): Refresh abstime before proceeding to LRU
eviction after an idle period, so that the next my_cond_timedwait()
will not return immediately on a stale deadline.

buf_LRU_scan_and_free_block(): Declare static.

buf_flush_LRU_list_batch(): In a run of blocks that
make_young_if_needed() moves, release and reacquire buf_pool.mutex
after every 512 scanned blocks, except on the first scanned block.
A move costs much less than an eviction, so the stride is longer than
the one of the eviction path.

buf_pool_invalidate(): Define in the same compilation unit with
buf_LRU_scan_and_free_block().

buf_pool.LRU_old_time_threshold: Replaces buf_LRU_old_threshold_ms.

PageConverter::run(): Renamed from fil_iterate(). In debug builds,
initialize and acquire a dummy exclusive latch on the block, so the
assertion in buf_page_t::invalidate() is satisfied; free the latch on
every path, not only on success.

AbstractCallback::m_zip_ssize: Replaces m_zip_size.

page_zip_des_t: Add calc_ssize()/zip_size() helpers, replacing the
zip_size<->ssize conversion duplicated across buf0buf.cc, buf0rea.cc
and row0import.cc.

innodb.buf_lru_scan_resistance: A new big test that checks that pages
of a table that is accessed in one burst are evicted, even if a sweep
reaches them only after innodb_old_blocks_time, that pages accessed
again after that time are promoted, and that pages read while the
buffer pool is being filled are not promoted later.

innodb_zip.n_blobs_700: A new test that stores 700 BLOB pointers on
one ROW_FORMAT=COMPRESSED page, within the 10-bit n_blobs field. The
index_page_splits counter confirms that the clustered index consists
of its root page alone, without depending on which pages remain in the
buffer pool.

Co-Authored-By: Alessandro Vetere <[email protected]>
Rucha Deodhar
MDEV-39049: Memory corruption & crash in check_key_in_list upon using
JSON_KEYS after modifying character set name/collation

Analysis:
Since the length of string is 0, accessing out of boundry memory,
leads to crash

Fix:
If string length is empty, return success.