Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
Oleksandr Byelkin
MDEV-41175 Fix stack-buffer-overflow in backup_log_ddl()

backup_log_ddl() sized its stack log buffer assuming each
identifier is at most ~40 bytes, but add_name_to_buffer() can
expand each identifier character to 5 bytes when re-encoding it
into my_charset_filename, so a RENAME TABLE with long, special-
character names overflowed the buffer (reported under ASAN).

Fixed by sizing the buffer to the true worst case per identifier,
and by making add_str_to_buffer() and its callers assert on an
explicit end-of-buffer pointer before writing.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Daniel Black
MDEV-41200 ANY_VALUE not marked BINLOG_STMT_UNSAFE_SYSTEM_FUNCTION

Without this marking the slave may differ from the master.
Rex Johnston
MDEV-41228 Test deep clones of Item_cache items

Deep clones of items are created in very few places, so most of the Item
clone code has no test coverage at all, although Parallel Query relies on it
heavily. One of the places where deep clones are created is the generation
of the key parts for a lookup into a split materialized table, in
TABLE::add_splitting_info_for_key_field().

An Item_cache reaches that place through the IN->EXISTS transformation.
Item_in_optimizer::fix_left() wraps the left expression of the predicate in
an Item_cache, and the equality injected into the subquery refers to it, so
the key field value being cloned is an Item_direct_ref over that cache. Two
more things are needed for the clone to happen: the grouping field the
equality matches has to be the first component of some index of the
underlying table, otherwise it is not among spl_opt_info->spl_fields and the
function returns before cloning, and the subquery must not be converted to a
semi-join. The latter is achieved by the shape of the query, a UNION in the
subquery, rather than by turning optimizer switches off, so that the plan is
the one a user gets with a default optimizer_switch.

We add Item::check_deep_copy(), which validates a clone against its original in
a debug build. It walks both item trees and reports, as notes, whether they
have the same shape with the same Item class at every node, and whether the
clone shares an Item object with the original. Sharing is what distinguishes
a shallow copy from a deep one: an item that is shallow by design, Item_field
for instance, still produces a separate object and only shares a Field, which
is not an Item.

Call it from TABLE::add_splitting_info_for_key_field() under the
"split_materialized_clones" debug flag, which additionally makes the
optimizer use a clone as the value of the generated key part. The clone then
has to work both in the condition pushed into the materialized table and as
the value looked up in the filled table. Note that the clone built for the
pushed condition cannot be reused for this, as it has already been made
dependent on the select that specifies the materialized table.

With the check in place, the clone of the cache turned out not to be a deep
one: every Item_cache_* class implemented deep_copy() as a plain shallow
copy. That is wrong beyond sharing the example item. A cache is filled by the
store()/cache_value() calls of the item that owns it, Item_in_optimizer here,
and nobody does that for a clone, so a clone that inherited the cached value
of the original kept returning that value for the rest of the query.

Implement Item_cache::deep_copy() once for all cache classes instead. The
clone is given an empty cache, so that it computes the value itself out of
the item the value is read from, and a copy of that item. The exception is an
example containing an aggregate or a window function: those are not clonable
yet, as a copy of one shares the per-execution data of the original and both
would free it, so such an example is shared and the check reports the clone
as not fully deep. Item_cache_row is not clonable at all now, as a copy of it
shared the values[] array of element caches with the original.

The test uses a single row in the outer table, because the answer to the same
query with more rows is wrong for an unrelated reason, MDEV-41251: a split
materialized table in a dependently executed subquery is never refilled when
the outer row changes.

We run our mtr test using the above like this
mtr --mysqld=--debug-dbug=+d,split_materialized_clones --suite=main \
  --parallel=auto
or similar.

Items of type Item_cache, the main Item class of interest in this commit,
are usually created by the type handler, through the call
Item::get_cache.  So another approach here is deep copy our Item_cache
here.

Deep copies of Item_caches can be tested using mtr like this
mtr --mysqld=--debug-dbug=+d,item_cache_clones --suite=main --parallel=auto
or similar.

Note that they can be combined, so anything Item_cache items used for
split materialized key generation will be a deep clone of a deep clone.

Another DBUG flag is "check_deep_copy".  This is used to check how many
items are deep cloned by producing a warning message.
Jaeheon Shim
Implement ANY_VALUE as a 'real' aggregate function through subclassing Item_sum_min_max
Dave Gosselin
MDEV-35845:  Propagate a constant into an IN predicate

SELECT * FROM t1 WHERE v IN ('a','b') AND v = 'b' kept both conjuncts
when v is a string column, while the equivalent form written with OR
was simplified to v = 'b'.

Two mechanisms can perform a rewrite.  Multiple equalities handle it
when check_simple_equality() builds an Item_equal, which it does only
if the field's charset allows constant propagation.  Up through 10.5 the
default character set was latin1 whose collation handler supports
constant propagation.  MDEV-19123 made utf8mb4 the default in 11.6, and
the utf8 collation handlers report that they do not support constant
propagation.

The other mechanism is propagate_cond_constants(), which rewrote
the OR form under every collation.  It descends through
change_cond_ref_to_const(), which returns on any node whose
eq_cmp_result() is COND_OK.  Item_func_in inherits that value, so the
IN predicate was skipped.

Implement an optimization in change_cond_ref_to_const() that replaces
the predicant of an IN predicate with the constant from an equality at
the same AND level.  The predicant is compared against every value of
the list, so the existing per-operand test from MDEV-7152 is applied
once for each of them.

Only a predicant whose arguments were all aggregated to one comparison
data type is replaced.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Jaeheon Shim
Add charset/collation/coercibility tests
Jaeheon Shim
Add prefix and suffix to any_value.test
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

In progress
Jaeheon Shim
Modify any_value test case to include null rows and empty table
Daniel Black
MDEV-39932 - remove test cases here. Already merged on rebase.
Rex Johnston
MDEV-41202 ONLY_FULL_GROUP_BY silently bypassed for outer-correlated fields inside aggregate

Item field->marker left unset when pushing items into the
relevant JOIN::non_agg_fields List.  field->marker is left
stale at MARKER_UNDEF_POS (0) rather than correctly set to
JOIN::select->cur_pos_in_select_list.
Any select list item not at position 0 is incorrectly evaluated
in setup_group(), where MODE_ONLY_FULL_GROUP_BY is processed.
Khaled Riyad
MDEV-40184 Heap-use-after-free in stored procedure when LEFT(CONCAT(...)) result is reused by LOCATE()

LEFT(), RIGHT() and SUBSTR() returned tmp_value pointing into the
buffer passed by the caller, without owning it. When a merged derived
column is used twice, e.g. LOCATE(a, x, a), the second evaluation of
the same item comes from Item_str_func::val_int(), which passes a
local buffer. tmp_value is re-pointed into that buffer, the buffer is
freed on return, and the first result becomes dangling.

Copy the result into tmp_value instead, as MDEV-32758 did for TRIM().
Jaeheon Shim
Modify existing group_by test to allow referencing outer field in correlated subquery as long as it is aggregated
Aleksey Midenkov
MDEV-40821 SIGSEGV in Window_funcs_sort::setup

Window_funcs_sort failed accessing win_func->window_spec as window
spec was not defined in the query. Usually this is checked by fixing
item win_func, but it was not done by setup_conds.

Actually, earlier check by vers_setup_conds() must fail on DELETE
HISTORY from non-versioned table, but TABLE_LIST for t1 has no
versioning conditions.

The cause was the parser assigning versioning conditions to wrong
table pointed by last_table() which was already switched to another
table from SYSTEM_TIME expression (cs1).

The fix assigns versioning conditions to the correct table stored to
correspondent_table by delete_single_table branch of the parser.

Same fix applied to delete_single_table_for_period (period_conditions),
with tests added for both cases.

Also DBUG_ASSERT in TABLE::vers_switch_partition() checked
delete_history on last_table(), which is also wrong for this case.
Checking the query_tables is correct, as delete_single_table puts
the DELETE target first there, and tables referenced later in the
statement are appended after it.
Jaeheon Shim
Recalculate digst hashes in perfschema tests
Oleksandr Byelkin
MDEV-41258 Crash in Item_func_nextval on prepared re-execution

Sequence prelocking entries were matched across re-executions by
db.str/table_name.str pointer identity, aliasing memory on the TABLE's
own mem_root. Once the owning table was closed and reopened, those
pointers went stale, leaving a dangling duplicate entry that still got
opened and crashed.
Fix: match entries by owning TABLE_LIST and position, deep-copied onto
the statement's own arena.

The relink was gated by has_prelocking_list, which stays true forever
once a trigger is involved, so it silently stopped running after the
first execution for such statements.
Fix: moved it into its own function, called unconditionally per table,
keeping the original gating, scoped to DML statements only.

A stale entry's linked_table pointer could survive its owner's open
being skipped, or a mid-statement open_tables() backoff
(close_tables_for_reopen()), leaving open_table()'s backlink write to
land on freed memory.
Fix: clear it in TABLE_LIST::reinit_before_use() and in
close_tables_for_reopen().

set_parameters() overwrote lex->default_used from each execution's own
bound parameters only, losing a literal DEFAULT written in the
statement text itself.
Fix: also record default_used at prepare time and OR it in.

A mid-execute_loop() reprepare() swaps in a freshly re-parsed lex,
losing "EXECUTE ... USING DEFAULT" already folded into default_used by
set_parameters() for this execution.
Fix: save that bit before reprepare() and OR it back in afterwards.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Jaeheon Shim
Update any_value tests involving window functions to use tables where the any_value column is functionally dependent on the partition by column. This is necessary as the iteration order when computing any_value on a window function is not guaranteed across executions.
bsrikanth-mariadb
MDEV-40837 Fix crash recording opt context with NULL character_set_results

Unlike character_set_client or collation_connection,
character_set_results can be set to NULL. When recording context for
a query in opt_context_store_replay.cc,
character_set_results->cs_name was accessed without first checking
whether character_set_results itself was NULL, causing a crash.

Fix Optimizer_context_recorder::dump_sql_script() to check
character_set_results for NULL before accessing cs_name, writing
"SET character_set_results=NULL;" to the replay script in that case
so replay reproduces the recorded session instead of silently
defaulting to the script's SET NAMES charset.

Added a test for the same, in opt_context_store_stats.test, that
also checks the recorded script contains the NULL SET statement.
Jaeheon Shim
MDEV-41138 Ensure ANY_VALUE not fully reserved
Jaeheon Shim
Implement copy_or_same
GoldenEmperor1177
MDEV-40453 Keep overflow edge case tests in place

What was wrong:
The former wrapping cases in the "Overflow edge cases" section of
mysqltest_expression_evaluation.test had been moved to the end of the
file, so the section no longer showed them.

How it is fixed:
Put the five cases back in their original places in that section. They
now die, so each is followed by an expected-error check that runs a
nested mysqltest instance. The duplicates at the end of the file are
removed.

How it is tested:
main.mysqltest_expression_evaluation, main.mysqltest and
main.mysqltest_string_functions pass, also with --ps-protocol and with
a UBSAN-built mariadb-test client. This changes only the test files.
Jaeheon Shim
Add additional test cases to any_value
Jaeheon Shim
Modify Item_field::fix_fields and Item_field::fix_outer_field to push to non_agg_fields only if BOTH in_sum_func and in_any_value are false
Jaeheon Shim
Allow in_any_value to bypass set_non_agg_field_used in fix_fields
Jaeheon Shim
Track ANY_VALUE nesting context in fix_fields
Aleksey Midenkov
MDEV-29483 Heap-use-after-free (Binary_string::copy()) with window functions

JOIN::make_aggr_tables_info(): a query with a window function buffers
join rows into a postjoin-aggr temporary table via copy_fields().
BLOB/TEXT fields were copied by Copy_field::do_field_eq(), a raw
memcpy of the packed record representation (length + pointer),
leaving the tmp table's field pointing at whatever storage backed
the source row. Once the underlying handler (e.g. InnoDB) frees or
reuses that storage on a later row fetch, any later read of the
buffered blob (e.g. Item_field::str_result() via result_field,
referenced from a WHERE/HAVING subquery predicate) dereferences a
dangling pointer.

The fix forces BLOB/TEXT fields to be deep-copied (do_save_blob())
for any query with window functions, the same way GROUP BY already
does via save_sum_fields.

save_sum_fields controls an unrelated decision (whether to defer
Item_sum computation instead of wiring an incremental result_field,
see the SUM_FUNC_ITEM handling in Create_tmp_table::add_fields())
and must not be conflated with blob-copy safety. Introduce a
separate save_blobs flag, threaded through
JOIN::create_postjoin_aggr_table(), create_tmp_table() and
Create_tmp_table, that independently controls the Copy_field::set()
deep-copy decision.

Create_tmp_table: add a save_blobs constructor parameter, plus an
overload preserving the historical combined behavior (save_blobs
same as save_sum_fields) for existing callers.
Oleksandr Byelkin
MDEV-41144 Assertion `scale <= precision' failed in decimal_bin_size

dynamic_column_decimal_read() derived intg/frac from an unbounded
var-uint and only checked scale>precision in int domain, so a crafted
value could still overflow decimal_bin_size()'s uint16 parameters,
tripping its scale<=precision assert (an out-of-bounds array read in
release builds).

Bound intg/frac against DECIMAL_MAX_POSSIBLE_PRECISION before use, and
reject the var-uint decoder's error signal, as the sibling
dynamic_column_string_read() already does. Validate the decoded
ulonglong before narrowing it to int, so a value wrapping around on
narrowing cannot bypass the bound.
bsrikanth-mariadb
MDEV-36986: Support tracing array of primitive types

Json_writer had two separate code paths: add_unquoted_str() for
numbers/bool/null, and add_escaped_str() which added the surrounding
quotes itself for strings. Single_line_formatting_helper, which
buffers consecutive array/object elements to decide if they fit on
one line, assumed only strings could ever be buffered and always
wrapped the flushed values in quotes. As a result, arrays of numbers
(e.g. "depends_on_map_bits", "rec_per_key") were incorrectly rendered
with their elements quoted as strings.

Unify both paths into add_escaped_quoted_str(): the caller now hands
over bytes that are already in their final on-the-wire form. String
escaping (json_escape_to_string) writes its own surrounding quotes,
while numbers/bool/null are passed through unquoted, so the one-line
helper just concatenates the buffered payloads on flush instead of
adding quotes itself. A DBUG_ASSERT in add_escaped_quoted_str() now
checks that every payload is already a quoted string or a bare
number/bool/null token, so a caller that violates the contract trips
an assertion in debug builds instead of silently producing invalid
JSON.

Also:
- Fix Json_writer_array::add(ulonglong)/(size_t), which went through
  add_ll() with a cast to longlong and corrupted large unsigned
  values (e.g. ULLONG_MAX); route them through add_ull() instead.
- Fix mysql-test/include/opt_context_schema.inc: "subquery_runs" was
  nested inside the preceding object instead of being a sibling
  member, and "rec_per_key" items are now declared as "number" to
  match the corrected output.
- Update recorded .result files for opt_trace, opt_context_*, and
  subselect_mat_analyze_json to reflect numbers/booleans no longer
  being quoted inside JSON arrays.
- Extend unittest/sql/my_json_writer-t.cc with coverage for arrays
  of primitives: plain integers, mixed types, sizes, multi-line
  arrays, values flushed before a nested object, strings that were
  already escaped by the one-line buffer, and invalid utf8mb4 input
  through add_str().
- Rewrite the "multi-line array of integers" test to actually
  overflow the one-line buffer (9 seven-digit numbers instead of 7),
  since the previous version fit on one line and never exercised the
  element-per-line flush path it claimed to test.
- Drop the now-redundant single-argument add_escaped_quoted_str()
  overload; all call sites already know their length.
- Update json_escape_to_string()'s doc comments (my_json_writer.h,
  sql_json_lib.h) to state that it quotes its output, not just
  escapes it.
Marko Mäkelä
fixup! 33ec54440d8b0ebd6fb4a2286c1167b851b1cdce
Dave Gosselin
MDEV-35747:  Wrong result from prepared TVC with parameter markers

The setup of column type information in table_value_constr::prepare()
was wrapped in an "if (!holders)" guard so that it runs only once per
statement.  However, the guard was too wide because it bound the
allocation of item holders (which should happen only once) to the
collection of type information (which should happen on each execution).

This leaves the TVC stuck with whatever placeholder type the parameter
had when the holders were first built, which may not match the type of
the next substitution.  A parameter marker has no type of its own until
a value is bound at EXECUTE time.  So both the TVC types and the
corresponding Item_type_holder instance in the SELECT item list must be
computed again on every EXECUTE.

Type holder allocation happens on the first call to the prepare()
function but that doesn't always coincide with a PREPARE.  It does for a
prepared statement whose table value constructor comes from the parser.
For a statement of a stored procedure, and for a table value constructor
that the conversion of an IN predicate into an IN subquery creates,
allocation happens instead on the first execution.  The corresponding
assertion allows the first execution and conventional execution as well
as PREPARE.

This patch separates the work done once per statement from the work done
on every execution as described above.

Whether the SELECT list of Item_type_holder instances has been built is
read from that list rather than from the holder array.  An error raised
while collecting the types leaves the array allocated and the list
empty, and the next call has to build the list.  Nullability starts over
on each collection so that it reflects the values of the current
execution.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
bsrikanth-mariadb
MDEV-36986: Support tracing array of primitive types

Json_writer had two separate code paths: add_unquoted_str() for
numbers/bool/null, and add_escaped_str() which added the surrounding
quotes itself for strings. Single_line_formatting_helper, which
buffers consecutive array/object elements to decide if they fit on
one line, assumed only strings could ever be buffered and always
wrapped the flushed values in quotes. As a result, arrays of numbers
(e.g. "depends_on_map_bits", "rec_per_key") were incorrectly rendered
with their elements quoted as strings.

Unify both paths into add_escaped_quoted_str(): the caller now hands
over bytes that are already in their final on-the-wire form. String
escaping (json_escape_to_string) writes its own surrounding quotes,
while numbers/bool/null are passed through unquoted, so the one-line
helper just concatenates the buffered payloads on flush instead of
adding quotes itself. A DBUG_ASSERT in add_escaped_quoted_str() now
checks that every payload is already a quoted string or a bare
number/bool/null token, so a caller that violates the contract trips
an assertion in debug builds instead of silently producing invalid
JSON.

Also:
- Fix Json_writer_array::add(ulonglong)/(size_t), which went through
  add_ll() with a cast to longlong and corrupted large unsigned
  values (e.g. ULLONG_MAX); route them through add_ull() instead.
- Fix mysql-test/include/opt_context_schema.inc: "subquery_runs" was
  nested inside the preceding object instead of being a sibling
  member, and "rec_per_key" items are now declared as "number" to
  match the corrected output.
- Update recorded .result files for opt_trace, opt_context_*, and
  subselect_mat_analyze_json to reflect numbers/booleans no longer
  being quoted inside JSON arrays.
- Extend unittest/sql/my_json_writer-t.cc with coverage for arrays
  of primitives: plain integers, mixed types, sizes, multi-line
  arrays, values flushed before a nested object, strings that were
  already escaped by the one-line buffer, and invalid utf8mb4 input
  through add_str().
- Rewrite the "multi-line array of integers" test to actually
  overflow the one-line buffer (9 seven-digit numbers instead of 7),
  since the previous version fit on one line and never exercised the
  element-per-line flush path it claimed to test.
- Drop the now-redundant single-argument add_escaped_quoted_str()
  overload; all call sites already know their length.
- Update json_escape_to_string()'s doc comments (my_json_writer.h,
  sql_json_lib.h) to state that it quotes its output, not just
  escapes it.
Marko Mäkelä
MDEV-41359 fix
Jaeheon Shim
Add additional any_value tests
- Tests without ONLY_FULL_GROUP_BY
- More NULL tests
- Window function tests
Jaeheon Shim
Add additional test cases to verify edge cases
Dave Gosselin
MDEV-41303:  rand() in a semi-join subquery is checked on outer rows

Converting an IN subquery to a semi-join moves its WHERE into the
parent WHERE.  A condition there such as rand(1) < 0.09 reads no table,
so it is attached to the last table of the join order that is outside
any materialized semi-join.  With SJ-Materialization it was checked once
for each outer row instead of once for each row of the subquery, and
the query returned a wrong count.

Do not convert a subquery to a semi-join when it contains a function
with a random result (UNCACHEABLE_RAND).  Derived tables already follow
this rule.  ROWNUM sets the same flag, so this replaces the check for
ROWNUM.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Jaeheon Shim
Undo COALESCE implementation of ANY_VALUE
GoldenEmperor1177
MDEV-40453 Reject overflowing mysqltest expressions

What was wrong:
The mysqltest expression evaluator did signed arithmetic without
checking that the result fits in long long. Boundary expressions
triggered undefined behavior or returned mathematically wrong wrapped
values, e.g. -9223372036854775808 / -1 gave -9223372036854775808, and
unsigned 18446744073709551615 + 1 gave 0.

How it is fixed:
Check unary minus, +, -, *, / and abs() before evaluating and die() with
an "Evaluation error: ... overflow" message when the result is not
representable in 64 bits. Unsigned +, - and * are checked the same way;
an unsigned difference that is negative but fits in long long is
returned as a signed value. A negative signed operand mixed with an
unsigned one is rejected instead of being silently converted. Valid
boundary results are kept: -9223372036854775808 % -1 = 0 and
-4611686018427387904 * 2 = -9223372036854775808. Unary minus on an
unsigned value still wraps, as conv(-~0, ...) in
main.mysqltest_string_functions relies on it.

How it is tested:
main.mysqltest_expression_evaluation: the former wrapping cases now
expect an error, and signed and unsigned overflow cases were added.
main.mysqltest, main.mysqltest_string_functions and the other tests
using $() pass, also with --ps-protocol.
Jaeheon Shim
Fix a minor indexing error in sql_yacc.yy, add ANY_VALUE_FUNC to is_aggr_sum_func()
Jaeheon Shim
Create Item_func_any_value and Create_func_any_value
Marko Mäkelä
MDEV-33966: buf_page_make_young() is a contention point

The buf_pool.LRU list needs to reasonably accurately reflect recently
accessed blocks, so that they will not be evicted prematurely.
Because the list is protected by buf_pool.mutex, it is not a good idea
to maintain the position on every page access.

Instead of maintaining the LRU position on each access, we will
decide on each access whether the block qualifies for promotion, and
record the decision in a "promote" flag in the block descriptor. The
buf_flush_page_cleaner() thread as well as some traversal of the
buf_pool.LRU list, all of which hold buf_pool.mutex anyway, will move
the flagged blocks to the "recently used" end of buf_pool.LRU.

A block qualifies only if it is accessed again at least
innodb_old_blocks_time after the reference point, which is its first
access and, after that, its most recent promotion. The age is measured
in whole seconds and checked on the access, not when a sweep reaches
the block, so that a block that a table scan accessed in one burst
will not be promoted, however long it stays in buf_pool.LRU_old. After
a promotion, the age is measured from the time the sweep moved the
block, which can be later than the access that qualified it. The rule
applies at any position in buf_pool.LRU, and the flag stays set until
a sweep moves the block. Thus, a block that qualified outside
buf_pool.LRU_old keeps the flag until a sweep reaches it, which is
usually after it has moved into buf_pool.LRU_old; before, such a
block was moved on any access when freed_page_clock showed that it was
no longer close to the "recently used" end.

buf_page_make_young_if_needed(), buf_page_make_young(),
buf_page_peek_if_too_old(), buf_page_peek_if_young(),
btr_cur_nonleaf_make_young(), buf_page_t::set_accessed(): Replaced by
buf_page_t::touch(), buf_page_t::touch_no_stamp() and
buf_page_t::make_young_if_needed().

buf_pool_t::freed_page_clock, buf_page_t::freed_page_clock: Remove.
This is no longer meaningful in the revised design.
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).FREE_PAGE_CLOCK now
always reports 0. Before the first eviction, buf_LRU_stat_update() now
records statistics intervals, and buf_LRU_evict_from_unzip_LRU() uses
its formula instead of assuming a disk-bound workload.

page_zip_des_t::state: An atomic 16-bit field that will include
the PROMOTE and OLD flags that would more logically belong
to buf_page_t. We maintain them here (along with some
ROW_FORMAT=COMPRESSED specific state that is protected by page
latches) in order to avoid race conditions and unnecessary
memory overhead. m_end moves out to its own field, shrinking the
packed word to 16 bits; n_blobs narrows from 12 bits to 10 to make
room for PROMOTE and OLD, still comfortably above the 744-column
maximum.

The helpers that set or clear one bit of page_zip_des_t::state always
perform the atomic read-modify-write, which is a full barrier on x86.
A caller that can find the bit already in the target state checks it
first: touch_no_stamp() for PROMOTE, buf_LRU_add_block() for OLD, and
page_zip_write_rec() for NONEMPTY. page_zip_des_t::copy_from() and
set_n_blobs_and_empty() skip the fetch_add() when the delta is 0.
page_zip_decompress_clust() counts the BLOB pointers locally and
invokes add_n_blobs() once.

page_zip_des_t::copy_from(): Replaces the copy constructor that
page_zip_copy_recs() invoked. It copies m_end, NONEMPTY and n_blobs
from the source page, but preserves the PROMOTE and OLD flags of the
destination block, because they describe the position of that block
in buf_pool.LRU. Copying OLD from the source would make
buf_pool.LRU_old_len inconsistent with the list. Like
set_n_blobs_and_empty(), it updates the state with a single
fetch_add(), which preserves any concurrent change of PROMOTE or OLD.

buf_page_t::init(): Takes the ROW_FORMAT=COMPRESSED shift size
(ssize) directly and clears the zip descriptor state itself, so
callers (buf_block_t::initialise(), buf_page_init_for_read()) no
longer need a separate page_zip_des_init()/page_zip_set_size() call.

buf_page_t::invalidate(): Replaces buf_block_modify_clock_inc().
Instead of maintaining a 64-bit counter, we will maintain one
comprising 32+16=48 bits, in modify_clock_low,modify_clock_high.
Worst case there will be exactly n<<48 calls to
buf_page_t::invalidate() before some operation such as
btr_pcur_t::restore_position() is executed. Such a count should
be extremely unlikely but not completely impossible. It is worth
noting that the DB_TRX_ID is only 48 bits, and each transaction
start and commit/rollback will consume an identifier.

buf_page_t::modify_clock(): Replaces the read access of
buf_block_t::modify_clock. Assert that the caller is holding
a page latch. Note: because invalidate() and modify_clock() are
protected with buf_pool.mutex or the buf_page_t::lock, there can
be no issue with regard to the atomicity of accessing the 48-bit field.

buf_page_t::relocate_from(): Copies a descriptor that is relocated
to this block, in buf_relocate() and buf_pool_t::shrink(), but keeps
the frame and the modify_clock() of this block. The copy constructor
does not copy the clock: it belongs to the block descriptor and must
only grow during its lifetime, as buf_block_t::modify_clock did.
Otherwise, the clock of this block could return to a value that a
btr_pcur_t saved before the block was freed, and
btr_pcur_t::restore_position() could reuse a record pointer although
the page was modified in another block in the meantime.

buf_page_t::access_time: Store the 16-bit buf_pool.access_clock
rather than a 32-bit millisecond ut_time_ms(). It wraps around every
18.2 hours; ages are computed as uint16_t(now - access_time), which
stays correct across that wrap. This avoids any alignment loss: the
adjacent fields modify_clock_low, modify_clock_high, access_time of
32+16+16 bits nicely add up to 64 bits. access_time is stamped on the
first access after the block was initialized, and on each promotion
by make_young_if_needed(), so that the age of a frequently promoted
block stays exact across the wrap.

buf_pool_t::access_clock: uint16_t(my_interval_timer() / 1000000000),
never 0 (see buf_pool_t::now()), refreshed about once per second by
buf_pool_t::refresh_clock() so that page accesses need not read the
system clock. srv_master_callback() refreshes it. From
buf_pool_t::create() until srv_master_timer is started, and for good
when srv_master_timer is not started (innodb_read_only,
innodb_force_recovery>=2, mariadb-backup), a separate
buf_pool_clock_timer refreshes it.

buf_pool_t::access_clock, buf_pool_t::LRU_old_threshold: Located in a
cache line of their own, because they are read on page accesses and
the adjacent buf_pool fields are frequently written.

buf_page_t::touch(): Stamp access_time on the first access, then
invoke touch_no_stamp(). Return whether this was not the first access,
as the result of buf_page_make_young_if_needed() used to be. A
buffer-fix is sufficient, as in MVCC undo page lookups.

buf_page_t::touch_no_stamp(): Set the PROMOTE flag if
innodb_old_blocks_time is 0, or if accessed_at() is at least that old,
at any position of the block in buf_pool.LRU. Because
buf_pool.access_clock advances once per refresh_clock(), about once per
second, a difference of threshold ticks can span less than
innodb_old_blocks_time; the difference must exceed the threshold, so
that a burst of accesses that crosses a refresh is not promoted. Like
btr_cur_nonleaf_make_young(), do not stamp access_time. Once PROMOTE is
set, later accesses only load the state and return.

buf_page_t::make_young_if_needed(): If PROMOTE is set, invoke
buf_page_t::make_young(). This part is inline, so that a sweep pays no
function call for a block that stays in place. It reads the state of
the block once: buf_pool.mutex, which the caller holds, protects OLD,
and only make_young() clears PROMOTE. The block need not be in
buf_pool.LRU_old: while buf_pool.LRU is shorter than
BUF_LRU_OLD_MIN_LEN, buf_pool.LRU_old does not exist, and blocks are
evicted without ever becoming old. The template parameter only_old
leaves a block that is not old in place; buf_do_flush_list_batch()
uses it, because it visits blocks in the order of oldest_modification,
and moving a block that is not old would not save it from eviction.
The template parameter count_not_young selects whether an old block
that has an accessed_at() stamp but is not flagged is counted in
buf_pool.stat.n_pages_not_made_young. This includes a block that had
no access since its last promotion, so the count is of sweep visits,
not of rejected accesses; only the eviction sweeps
buf_LRU_free_from_common_LRU_list() and buf_flush_LRU_list_batch() do
this, and a block that a sweep leaves in the list can be counted again
by a later sweep. As before, buf_pool.stat.n_pages_made_young counts
only the moves of old blocks. Because a sweep moves the block, not the
access, and counts the blocks that it does not move,
Innodb_buffer_pool_pages_made_young and
Innodb_buffer_pool_pages_made_not_young can differ from before for the
same workload.

buf_page_t::make_young(): Clear PROMOTE, stamp access_time, and move
the block to the "recently used" end of buf_pool.LRU. The block can be
read-fixed, because buf_pool_t::unzip() copies PROMOTE and OLD from the
compressed-only descriptor and releases buf_pool.mutex during
buf_zip_decompress(). Unlike buf_page_make_young(), we do not skip such
a block: a page access no longer acquires buf_pool.mutex, and the
sweeps hold buf_pool.mutex but no buffer-fix on the block.

buf_LRU_remove_block(): When buf_pool.LRU becomes too short for
buf_pool.LRU_old, clear the OLD flags from buf_pool.LRU_old to the end
of the list, which holds all the old blocks, not the flags of every
block. While buf_pool.LRU_old does not exist, no block in buf_pool.LRU
is old, so on later removals the loop does nothing.

buf_LRU_add_block(): Write the OLD flag only if it changes. In a list
that is too short for buf_pool.LRU_old, assert that buf_pool.LRU_old
is not defined and that the block is not old, instead of writing the
flag.

innodb_old_blocks_time_update(): New sysvar update callback,
replacing a NULL one, that calls buf_pool_t::set_old_threshold_ms(),
so that SET GLOBAL innodb_old_blocks_time also updates the
LRU_old_threshold in seconds that page accesses read. The threshold
is clamped to 65534 seconds (UINT16_MAX - 1), because the age
uint16_t(now - access_time) cannot exceed 65535 and a larger threshold
would prevent any promotion. The sysvar itself still accepts up to
UINT_MAX32 milliseconds.

buf_pool_t::set_old_threshold_ms(): Round the millisecond threshold
up, not down, to the nearest second. innodb_old_blocks_time is
documented and accepted in milliseconds; flooring instead of ceiling
would make any configured value from 1 to 999 silently behave as 0
(disabled).

buf_page_t::is_accessed(): Renamed accessed_at(), to stop reading as
a boolean. It returns the access_time stamp of the first access or
of the last promotion, in seconds;
INFORMATION_SCHEMA.INNODB_BUFFER_PAGE(_LRU).ACCESS_TIME now reflects
that.

buf_read_ahead_random(): A page now qualifies once accessed_at()
holds, together with either zip.is_promote() or !zip.old().

buf_read_ahead_linear(): Compare access_time stamps as a signed
16-bit difference, not raw unsigned, so the monotonic-access check
stays correct across the access_time wrap. The resolution of the
stamps is 1 second instead of 1 millisecond.

buf_flush_page_cleaner(): Refresh abstime before proceeding to LRU
eviction after an idle period, so that the next my_cond_timedwait()
will not return immediately on a stale deadline.

buf_LRU_scan_and_free_block(): Declare static.

buf_flush_LRU_list_batch(): In a run of blocks that
make_young_if_needed() moves, release and reacquire buf_pool.mutex
after every 512 scanned blocks, except on the first scanned block.
A move costs much less than an eviction, so the stride is longer than
the one of the eviction path.

buf_pool_invalidate(): Define in the same compilation unit with
buf_LRU_scan_and_free_block().

buf_pool.LRU_old_time_threshold: Replaces buf_LRU_old_threshold_ms.

PageConverter::run(): Renamed from fil_iterate(). In debug builds,
initialize and acquire a dummy exclusive latch on the block, so the
assertion in buf_page_t::invalidate() is satisfied; free the latch on
every path, not only on success.

AbstractCallback::m_zip_ssize: Replaces m_zip_size.

page_zip_des_t: Add calc_ssize()/zip_size() helpers, replacing the
zip_size<->ssize conversion duplicated across buf0buf.cc, buf0rea.cc
and row0import.cc.

innodb.buf_lru_scan_resistance: A new big test that checks that pages
of a table that is accessed in one burst are evicted, even if a sweep
reaches them only after innodb_old_blocks_time, that pages accessed
again after that time are promoted, and that pages read while the
buffer pool is being filled are not promoted later.

innodb_zip.n_blobs_700: A new test that stores 700 BLOB pointers on
one ROW_FORMAT=COMPRESSED page, within the 10-bit n_blobs field. The
index_page_splits counter confirms that the clustered index consists
of its root page alone, without depending on which pages remain in the
buffer pool.

Co-Authored-By: Alessandro Vetere <[email protected]>