Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations (latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE) for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collation" error can be avoided in a comparison
operator. We can choose any of the sides as the operation effective
collation - the result will be equal.

This optimization is not applied when both sides have an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison is
already illegal when the character sets are the same, so for
consistency it stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality on this repertoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st

  It returns the tailoring on the given repertoire for the given collation.

  If cs1->cset->tailoring(cs1, some_repertoire) returns {0,0},
  it means illegal mix optimization cannot be used for this collation
  on the given repertoire.

  If these calls:
    tr1= cs1->cset->tailoring(cs1, some_repertoire);
    tr2= cs2->cset->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.ptr==tr2.ptr,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Adding a new method DTCollation::aggregate_by_tailoring().

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose a collation
  of one of the sides (even if they are compatible on the given repertoire).

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits
  rather than a single bit, the way how to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (!(repertoire & ~MY_REPERTOIRE_ASCII))

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII repertoire.

  * MY_CS_IDENT_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT repertoire
    (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII letters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_BINARY_CI and MY_CS_IDENT_BINARY_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as returned by:
    cs->cset->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests.

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';
ParadoxV5
MDEV-38849 slave_connections_needed_for_purge prevents independent machine from purging binary logs

`@@slave_connections_needed_for_purge`’s default of `1` ensures binary
log availability on replication masters, but is not a sensible default
suitable for all scenarios, especially for long-term slave servers and
standalone (not in a replication setup) servers.
The outcome was that standalone server users were confused why automatic
binlog purging does not work.

This commit changes this default to `0`, which is suitable for both
standalone and (when backed by prompt failure recovery)
replication setups.
`0` also more closely matches the behaviour before MDEV-31404,
which added this variable, out of the box.

This commit also adds a one-time replication warning when registering a
slave, but `@@slave_connections_needed_for_purge` is left unchanged.
Rather than enforcing a defence with an unsensible default, this
reminder will bring awareness of the risk of automatic binlog purging.

This commit also cleans up Galera and MTR workarounds to the
introduction of the `@@slave_connections_needed_for_purge=1` default.
Daniel Black
MDEV-41156: ASAN use-after-poison in JSON_CONTAINS_PATH

The json_depth_array as allocated for 32 elements but accessed
it as though there wasn't a limit.

Called mem_root_dynamic_array_resize_and_get_val to ensure that
the number of elements was presented, and error JE_EOS (out of
space) if allocation exceeded and return its pointer.

With the array resizing called early mem_root_dynamic_array_resize_and_set_val
isn't required.

Removed unused 'value' variable, as value_ptr was always valid.

Reported by: David Korczynski of Ada Logics
DerZc
MDEV-40689 Wrong result: BIT_AND/BIT_OR/BIT_XOR in WINDOW functions over a frame containing NULL

BIT_AND, BIT_OR, and BIT_XOR window functions can return incorrect
values as a sliding frame moves past NULL input rows.

Adding a NULL argument leaves the bit-aggregate state unchanged, but
removing that row unconditionally calls remove_as_window() with
val_int()'s value. The removal path therefore changes state for a row
that never contributed to the aggregate.

Evaluate the departing window argument once and retain its unsigned
value. Call remove_as_window() only when the evaluated argument is
non-NULL. Leave the existing incremental window algorithm and non-window
aggregation path in place.

The regression checks all three bit aggregates over ROWS BETWEEN 1
PRECEDING AND CURRENT ROW with interleaved NULL and non-NULL values,
including removal of a real zero and restoration of the neutral values
after the frame becomes all-NULL.

Bug report: https://jira.mariadb.org/browse/MDEV-40689
Khaled Riyad
MDEV-38004 Double free on re-execution of prepared aggregate function

Between two executions of a prepared statement, Item_sp::cleanup() freed
the stored function's memory root but left the arena's free list pointing
into it. That list is filled when the function call returns and the active
arena is restored.

On the next execution Item_sum_sp::clear() walked the stale list and
destroyed already freed items.

Free the items before the memory they live in.  The sequence is now
in Item_sp::free_call_ctx(), so the three places that repeat it
cannot drift apart again.
Dave Gosselin
MDEV-41211 Remove unused index_read_idx() from both Federated engines

ha_federated::index_read_idx() and ha_federatedx::index_read_idx()
have no callers and do not override a handler method.  Remove them,
and remove the sentence in the comment on each index_read() that says
index_read() calls it.  Comments that describe index_read_idx() now
name index_read() or index_read_idx_map(), whichever does that read.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Dave Gosselin
MDEV-33616:  Take the read lock many times in perfschema.func_mutex

The wait timer can have a granularity coarser than the time an
uncontended read lock is held, so the recorded duration of one lock can
be zero, which reads back as NULL.  This can cause the test to fail with
a false negative.

Take the lock twenty more times at each measurement point, with the
extra statements silent so the recorded result does not change.  The
mutex part of the test already works this way, since one SELECT
produces ten THR_LOCK::mutex events.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
KhaledR57
MDEV-32383 Server crashes in Item_func_match::init_search on 2nd execution of PS

Item_func_match::master points at an equal MATCH item that owns the
shared ft_handler. setup_ftfuncs() sets it only when it is still unset,
and cleanup() never reset it.

A mergeable view is merged once. On re-execution mysql_derived_prepare()
returns early because TABLE_LIST::merged is set, so the view's own MATCH
item is never re-fixed and keeps the NULL table left by cleanup().
init_ftfuncs() skips unfixed items in ftfunc_list, but init_search()
follows master without that check and dereferenced the NULL table.

Reset master in cleanup(), after the ft_handler ownership check that
reads it, so the link is rebuilt from scratch on every execution.
Dave Gosselin
MDEV-33616:  Skip the redo log upgrade tests without sparse file support

innodb.log_upgrade and innodb.log_upgrade_101_flags build 8GB redo log
files by seeking past the end of an empty file and writing a single
byte.  That needs a filesystem which leaves the skipped range
unallocated.  HFS on macOS allocates every block of it instead, so the
write fails with ENOSPC and the test reports a perl failure.

include/have_sparse_files.inc probes a directory the caller names,
writing one byte 64MB into an empty file there and comparing the
allocated block count against that offset.
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations
(latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE)
for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collations" error can be avoided in a comparison
operator. We can choose any of the sides as the operation effective
collation - the result will be equal.

This optimization is not applied when at least one side has an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison are
already illegal when the character sets are the same, so for
consistency this stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality
  on these repertoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st.

  It returns the tailoring on the given repertoire for the given collation.

  If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0},
  it means the illegal mix optimization cannot be used
  for this collation on the given repertoire.

  If these calls:
    tr1= cs1->coll->tailoring(cs1, some_repertoire);
    tr2= cs2->coll->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.str==tr2.str,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Adding a new method DTCollation::aggregate_by_tailoring().
  Like the other successful branches of DTCollation::aggregate(),
  it makes the resulting repertoire cover both sides.

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose
  a collation of one of the sides (even if they are compatible
  on the given repertoire). This also applies to ExtractValue()
  and UpdateXML(), which now aggregate their arguments without
  this flag.

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits
  rather than a single bit, the way to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (!(repertoire & ~MY_REPERTOIRE_ASCII))

- Repertoire of some expressions was adjusted to the new meaning:
  * Item_null now has MY_REPERTOIRE_NONE (was ASCII).
  * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_ALL
    (was EXTENDED).
  * Lex_string_with_metadata_st::repertoire(cs) now scans the string
    contents to detect the actual repertoire.
  * HEX() has a narrower repertoire.
  * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD and GROUP_CONCAT add
    the repertoire of the extra characters they put into the result
    (quotes, separators, padding).

- The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL),
  so my_charset_is_ascii_based() is true for tis620.
  This changes the result of subselect_extra_no_semijoin
  from an error to success.

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII
    repertoire.

  * MY_CS_IDENT_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT
    repertoire (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII characters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_BINARY_CI and MY_CS_IDENT_BINARY_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as returned by:
    cs->coll->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests.

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';

- Adding tests to ctype_cp932.test checking that ExtractValue() and
  UpdateXML() still raise "Illegal mix of collations".
Sergei Petrunia
Add comment about Create_tmp_table::m_group
Khaled Riyad
MDEV-38861 heap-use-after-free in Prepared_statement::execute()

DROP PROCEDURE and CREATE OR REPLACE PROCEDURE executed from inside the
routine itself removed it from the SP cache. sp_head::destroy() then freed
the memory root that the running sp_head, its LEX and its instructions
live in, and the caller kept using them.

Skip the removal while the routine is being executed. sp_cache_invalidate()
above has already bumped the cache version, so the stale entry is removed by
the next lookup, after IS_INVOKED has been cleared.
Aleksey Midenkov
MDEV-38758 vcol.vcol_partition_innodb fails on buildbot

TABLE_ROWS in INFORMATION_SCHEMA.PARTITIONS for InnoDB is
dict_table_t::stat_n_rows. Each INSERT puts the record into the
clustered index and only afterwards increments stat_n_rows without
any latch. The second INSERT into a fresh partition schedules a
background persistent stats recalc, which is not throttled because
stats_last_recalc is still 0. If the recalc scans the leaf page
after the next INSERT has added its record but before that INSERT
increments stat_n_rows, the recalc stores the exact count and the
increment then adds one more, so p2 shows 4 rows instead of 3.

The fix disables innodb_stats_auto_recalc for the duration of
inc/vcol_partition.inc in the InnoDB variant of the test, so
stat_n_rows is maintained only by the per-row increments and is
exact.
Dave Gosselin
MDEV-41303:  rand() in a semi-join subquery is checked on outer rows

Do not merge a subquery into its parent as a semi-join when it has
the UNCACHEABLE_RAND flag, which RAND() and ROWNUM set.  Derived
tables already follow this rule.  ROWNUM sets the same flag, so this
patch replaces the check for ROWNUM with a check for UNCACHEABLE_RAND.

Previously, converting an IN subquery to a semi-join moved its WHERE
into the parent WHERE.  A condition there such as rand(1) < 0.09
doesn't rely on any columns, so it is attached to the last table of
the join order that is outside any materialized semi-join.  With
SJ-Materialization it was checked once for each outer row instead of
once for each row of the subquery, and the query returned a wrong
count.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Vladislav Vaintroub
MDEV-39533 Resolve reparse points and enforce MY_NOSYMLINKS on Windows

On Windows, my_realpath() only called GetFullPathName(), does not resolve
symlinks, junctions or mount points, unlike POSIX realpath(). At the
same time, my_open() and my_delete() ignored MY_NOSYMLINKS entirely,
so the symlink-attack protection used for MyISAM/Aria's DATA
DIRECTORY/INDEX DIRECTORY (mi_open()/ma_open(),
my_handler_delete_with_symlink()) was silently absent on Windows.

Fix my_realpath() to actually resolve reparse points: open the
path with CreateFile(), which follows them, and read back the handle's
fully resolved path with GetFinalPathNameByHandle(). As a result,
a missing path now correctly returns 1/ENOENT on Windows too, matching
Linux's realpath()-based behavior, instead of always returning 0.

Make my_open() and my_delete() honor MY_NOSYMLINKS on Windows.
Windows has no per-path-component O_NOFOLLOW equivalent, so instead
this mirrors the HAVE_REALPATH branch of the POSIX
NOSYMLINK_FUNCTION_BODY macro: the caller-supplied name (expected to
already be my_realpath()-resolved) is compared against the actually
opened handle's resolved path, and rejected with ENOTDIR -- the same
errno POSIX uses for this exact "not already canonical" condition --
on a mismatch, whether caused by a TOCTOU symlink swap or by the name
never having been fully resolved to begin with.

Known limitation: GetFinalPathNameByHandle(FILE_NAME_NORMALIZED), used
by both my_realpath() and MY_NOSYMLINKS verification, can fail on some
SMB shares (an intermediate directory denying list/read access while
still allowing traverse). When that happens, both fall back to the
weaker FILE_NAME_OPENED query: my_realpath() still succeeds, but
MY_NOSYMLINKS verification is weaker, since FILE_NAME_OPENED may not
fully resolve reparse points. A one-time warning naming the affected
path is raised the first time this happens.

As part of enforcing MY_NOSYMLINKS for my_delete(), my_win_unlink()
(formerly in my_delete.c) is rewritten and moved to my_winfile.cc
It opens file once, and verifies no symlinks, before removing it.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Dave Gosselin
MDEV-41211 Federated multi-table DELETE keeps a const table row

A multi-table DELETE on a FEDERATED or FederatedX table went wrong
when a primary key lookup made the target a const table.  The
optimizer reads a const table's row through index_read_idx_map(),
whose default implementation ends the index scan and frees the result
set.  The server asks for the row's position later, during execution,
so the saved position was empty.  FederatedX skipped the row and
FEDERATED crashed in rnd_pos().

Both engines now override index_read_idx_map() so that the lookup
leaves its result set open, as index_read() does.  position() then
records a valid position, and the result set is freed at the end of
the statement.  FEDERATED's reset() now also clears stored_result,
which otherwise pointed at a freed result set and was freed again
when the table was closed.

Tests include a multitable UPDATE with a const target, which crashed
both engines before the fix.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations (latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE) for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collation" error can be avoided in a comparison
operator. We can choose any of the sides as the operation effective
collation - the result will be equal.

This optimization is not applied when both sides have an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison is
already illegal when the character sets are the same, so for
consistency it stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPEROIRE_XXX flags set.

  This patch implements detecting tailoring equality on this reperoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st

  It returns the tailoring on the given repertoire for the given collation.

  If cs1->cset->tailoring(cs1, some_repertoire) returns {0,0},
  it means illegal mix optimization cannot be used for this collation
  on the given repertoire.

  If these calls:
    tr1= cs1->cset->tailoring(cs1, some_repertoire);
    tr2= cs2->cset->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.ptr==tr2.ptr,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Adding a new method DTCollation::aggregate_by_tailoring().

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose a collation
  of one of the sides (even if they are compatible on the given repertoire).

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits
  rather than a single bit, the way how to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (!(repertoire & ~MY_REPERTOIRE_ASCII))

- New flags were added int for CHARSET_INFO::state
  * MY_CS_ASCII_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation does not has no irregularities
    on the ASCII range.

  * MY_CS_IDENT_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has not irregularities
    on the IDENT subrange only (but can have irregularities say on
    punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII letters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_BINARY_CI and MY_CS_IDENT_BINARY_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as returned by:
    cs->cset->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests.

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.

- Adding a number of MTR tests in the form of the originally
  reported stript for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';
Rucha Deodhar
MDEV-41181: ASAN heap-buffer-overflow after SELECT JSON_SCHEMA_VALID
Georg Richter
MDEV-40146 vio_gencert function doesn't set serial number

this apparently breaks RFC 5280, and makes python cryptography module
unhappy.
Vladislav Vaintroub
build mysqlservices without an embedded CRT requirement

mysqlservices only exposes a thin C API, no CRT state crosses it, so
don't force whatever CRT/config built the server onto a plugin linking
it. Without /Zl, a plugin built in a config with no matching installed
mysqlservices variant (CMake silently substitutes one - verified with
a toy project) gets an ignorable but noisy LNK4098 warning.

Assisted-by: Claude:claude-5-sonnet
Dave Gosselin
MDEV-31180:  MyISAMMRG Crash on UPDATE of an updateable VIEW

Attach the children of a MERGE table once per statement, and keep the
value of pos_in_table_list for a MERGE table on subsequent executions
of a prepared statement.
Mohammad Tafzeel Shams
MDEV-41242 : Fix resource leaks on InnoDB/mariabackup error paths found by Infer

Several error-handling paths returned without releasing a resource
already acquired earlier in the function, or checked the wrong handle
entirely, risking use of an unopened handle.

Changes:
- SysTablespace::read_lsn_and_check_flags(): close the datafile handle
  on header-validation failure.
- xb_process_datadir(): check the freshly opened `dir` handle instead
  of the stale `dbdir`, fixing a handle leak and a possible use of an
  unopened directory handle.
- wsrep.cc / xb_load_list_file(): close file handles before die(), and
  null-check fopen() results in wsrep.cc.
- datadir_iter_new(): free datadir_path and destroy the mutex on the
  os_file_opendir() failure path.
Brad Smith
crc32c: check elf_aux_info() return value in ppc64 probe

elf_aux_info(3) leaves the output buffer unmodified on failure, so
ignoring the return value could test an uninitialized cpufeatures and
wrongly enable the POWER8 vector-crypto path.

Treat failure as "no features" so the probe falls back to the generic
implementation.
Aleksey Midenkov
MDEV-29566 parts.partition_special_innodb unstable DROP TABLE fix

The test sporadically fails with ER_LOCK_WAIT_TIMEOUT on DROP TABLE t1
of the Bug#53676 test case (see MDEV-29566).

  mysqltest: At line 119: query 'DROP TABLE t1' failed:
  ER_LOCK_WAIT_TIMEOUT (1205): Lock wait timeout exceeded; try restarting transaction

The DROP TABLE is executed by con2 with lock_wait_timeout=0. The INSERT
puts two rows into each partition, which exceeds the auto recalc
threshold of persistent statistics, so the partitions are added to the
recalc pool. The statistics background thread opens the partitions with
a shared MDL on t1, and if it does so at the moment of DROP TABLE, the
MDL request fails instantly.

The fix creates the table with STATS_PERSISTENT=0. The test case is not
about statistics; this also avoids locking mysql.innodb_table_stats and
mysql.innodb_index_stats in ha_innobase::delete_table() when the failed
ALTER TABLE ... ADD PARTITION rolls back the new partitions.
Vladislav Vaintroub
MDEV-39533 Resolve reparse points and enforce MY_NOSYMLINKS on Windows

On Windows, my_realpath() only called GetFullPathName(), does not resolve
symlinks, junctions or mount points, unlike POSIX realpath(). At the
same time, my_open() and my_delete() ignored MY_NOSYMLINKS entirely,
so the symlink-attack protection used for MyISAM/Aria's DATA
DIRECTORY/INDEX DIRECTORY (mi_open()/ma_open(),
my_handler_delete_with_symlink()) was silently absent on Windows.

Fix my_realpath() to actually resolve reparse points: open the
path with CreateFile(), which follows them, and read back the handle's
fully resolved path with GetFinalPathNameByHandle(). As a result,
a missing path now correctly returns 1/ENOENT on Windows too, matching
Linux's realpath()-based behavior, instead of always returning 0.

Make my_open() and my_delete() honor MY_NOSYMLINKS on Windows.
Windows has no per-path-component O_NOFOLLOW equivalent, so instead
this mirrors the HAVE_REALPATH branch of the POSIX
NOSYMLINK_FUNCTION_BODY macro: the caller-supplied name (expected to
already be my_realpath()-resolved) is compared against the actually
opened handle's resolved path, and rejected with ENOTDIR -- the same
errno POSIX uses for this exact "not already canonical" condition --
on a mismatch, whether caused by a TOCTOU symlink swap or by the name
never having been fully resolved to begin with.

Known limitation: GetFinalPathNameByHandle(FILE_NAME_NORMALIZED), used
by both my_realpath() and MY_NOSYMLINKS verification, can fail on some
SMB shares (an intermediate directory denying list/read access while
still allowing traverse). When that happens, both fall back to the
weaker FILE_NAME_OPENED query: my_realpath() still succeeds, but
MY_NOSYMLINKS verification is weaker, since FILE_NAME_OPENED may not
fully resolve reparse points. A one-time warning naming the affected
path is raised the first time this happens.

As part of enforcing MY_NOSYMLINKS for my_delete(), my_win_unlink()
(formerly in my_delete.c) is rewritten and moved to my_winfile.cc
It opens file once, and verifies no symlinks, before removing it.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
ParadoxV5
Merge branch '11.4' into MDEV-38849
Sergei Golubchik
fix the build for -G "Ninja Multi-Config"
Dave Gosselin
MDEV-28509:  Dereferenced null pointer of type 'struct JOIN_TAB' in add_key_field

setup_group no longer writes Item::marker.  The ONLY_FULL_GROUP_BY
check now tests GROUP BY membership by walking the GROUP BY list.

A query that defines a WINDOW but never refers to it could crash in
add_key_field, for example

  WITH cte AS (SELECT i FROM (SELECT i FROM t1 GROUP BY i) dt
              WINDOW w AS (PARTITION BY i))
  SELECT a.i FROM cte a JOIN cte b ON a.i=b.i WHERE a.i != 5;

A query that defines a WINDOW goes through setup_group, which set
marker to MARKER_UNDEF_POS (-1) on each GROUP BY expression so that
the ONLY_FULL_GROUP_BY check could skip it.  Other code reads marker
as a set of flag bits (-1 sets all bits).
Item_direct_view_ref::grouping_field_transformer_for_where then took
the ref as flagged for substitution and followed a path that ends in
the crash.

The ONLY_FULL_GROUP_BY check was the only reader of that value, so
MARKER_UNDEF_POS is removed.  The necessary check is local to the
setup_group function.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations (latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE) for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collation" error can be avoided in a comparison
operator. We can choose any of the sides as the operation effective
collation - the result will be equal.

This optimization is not applied when both sides have an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison is
already illegal when the character sets are the same, so for
consistency it stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPEROIRE_XXX flags set.

  This patch implements detecting tailoring equality on this reperoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st

  It returns the tailoring on the given repertoire for the given collation.

  If cs1->cset->tailoring(cs1, some_repertoire) returns {0,0},
  it means illegal mix optimization cannot be used for this collation
  on the given repertoire.

  If these calls:
    tr1= cs1->cset->tailoring(cs1, some_repertoire);
    tr2= cs2->cset->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.ptr==tr2.ptr,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Adding a new method DTCollation::aggregate_by_repertoire().

- Adding a new flag MY_COLL_ALLOW_BY_REPERTOIRE.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_REPERTOIRE.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose a collation
  of one of the sides (even if they are compatible on the given repertoire).

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits
  rather than a single bit, the way how to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (!(repertoire & ~MY_REPERTOIRE_ASCII))

- New flags were added int for CHARSET_INFO::state
  * MY_CS_ASCII_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation does not has no irregularities
    on the ASCII range.

  * MY_CS_IDENT_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has not irregularities
    on the IDENT subrange only (but can have irregularities say on
    punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII letters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_BINARY_CI and MY_CS_IDENT_BINARY_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as returned by:
    cs->cset->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests.

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.

- Adding a number of MTR tests in the form of the originally
  reported stript for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';
KhaledR57
MDEV-32383 Server crashes in Item_func_match::init_search on 2nd execution of PS

Item_func_match::master points at an equal MATCH item that owns the
shared ft_handler. setup_ftfuncs() sets it only when it is still unset,
and cleanup() never reset it.

A mergeable view is merged once. On re-execution mysql_derived_prepare()
returns early because TABLE_LIST::merged is set, so the view's own MATCH
item is never re-fixed and keeps the NULL table left by cleanup().
init_ftfuncs() skips unfixed items in ftfunc_list, but init_search()
follows master without that check and dereferenced the NULL table.

Reset master in cleanup(), after the ft_handler ownership check that
reads it, so the link is rebuilt from scratch on every execution.
Vladislav Vaintroub
MDEV-39533 Resolve reparse points and enforce MY_NOSYMLINKS on Windows

On Windows, my_realpath() only called GetFullPathName(), does not resolve
symlinks, junctions or mount points, unlike POSIX realpath(). At the
same time, my_open() and my_delete() ignored MY_NOSYMLINKS entirely,
so the symlink-attack protection used for MyISAM/Aria's DATA
DIRECTORY/INDEX DIRECTORY (mi_open()/ma_open(),
my_handler_delete_with_symlink()) was silently absent on Windows.

Fix my_realpath() to actually resolve reparse points: open the
path with CreateFile(), which follows them, and read back the handle's
fully resolved path with GetFinalPathNameByHandle(). As a result,
a missing path now correctly returns 1/ENOENT on Windows too, matching
Linux's realpath()-based behavior, instead of always returning 0.

Make my_open() and my_delete() honor MY_NOSYMLINKS on Windows.
Windows has no per-path-component O_NOFOLLOW equivalent, so instead
this mirrors the HAVE_REALPATH branch of the POSIX
NOSYMLINK_FUNCTION_BODY macro: the caller-supplied name (expected to
already be my_realpath()-resolved) is compared against the actually
opened handle's resolved path, and rejected with ENOTDIR -- the same
errno POSIX uses for this exact "not already canonical" condition --
on a mismatch, whether caused by a TOCTOU symlink swap or by the name
never having been fully resolved to begin with.

As part of enforcing MY_NOSYMLINKS for my_delete(), my_win_unlink()
(formerly in my_delete.c) is rewritten and moved to my_winfile.cc
It opens file once, and verifies no symlinks, before removing it.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations
(latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE)
for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collations" error can be avoided in a comparison
operator. We can choose any of the sides as the operation effective
collation - the result will be equal.

This optimization is not applied when at least one side has an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison are
already illegal when the character sets are the same, so for
consistency this stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality
  on these repertoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st.

  It returns the tailoring on the given repertoire for the given collation.

  If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0},
  it means the illegal mix optimization cannot be used
  for this collation on the given repertoire.

  If these calls:
    tr1= cs1->coll->tailoring(cs1, some_repertoire);
    tr2= cs2->coll->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.str==tr2.str,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Adding a new method DTCollation::aggregate_by_tailoring().
  Like the other successful branches of DTCollation::aggregate(),
  it makes the resulting repertoire cover both sides.

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose
  a collation of one of the sides (even if they are compatible
  on the given repertoire). This also applies to ExtractValue()
  and UpdateXML(), which now aggregate their arguments without
  this flag.

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits
  rather than a single bit, the way to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (!(repertoire & ~MY_REPERTOIRE_ASCII))

- Repertoire of some expressions was adjusted to the new meaning:
  * Item_null now has MY_REPERTOIRE_NONE (was ASCII).
  * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_ALL
    (was EXTENDED).
  * Lex_string_with_metadata_st::repertoire(cs) now scans the string
    contents to detect the actual repertoire.
  * HEX() has a narrower repertoire.
  * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD and GROUP_CONCAT add
    the repertoire of the extra characters they put into the result
    (quotes, separators, padding).

- The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL),
  so my_charset_is_ascii_based() is true for tis620.
  This changes the result of subselect_extra_no_semijoin
  from an error to success.

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII
    repertoire.

  * MY_CS_IDENT_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT
    repertoire (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII characters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_BINARY_CI and MY_CS_IDENT_BINARY_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as returned by:
    cs->coll->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests.

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';

- Adding tests to ctype_cp932.test checking that ExtractValue() and
  UpdateXML() still raise "Illegal mix of collations".
Sergei Golubchik
MDEV-40608 MariaDB-devel is incomplete for plugins

This works on Linux and on Windows, with rpm/deb/tar.gz/zip
installations.

For rpm/deb it just works, for tar.gz/zip there is no
standard location, so one needs to configure plugin with

  -DCMAKE_PREFIX_PATH=/pah/to/mariadb/basedir

after that, `cmake --install .` works too, installing in the same
basedir.

`cmake --build . --target package` works, creating rpm/deb/targz/zip
depending on whether it's Linux or Windows and whether -DRPM or -DDEB
was specified.

* create and install mariadb-plugin-config.cmake
* for now it only supports one plugin per project, error out
  if there are many
* deb: move all headers that plugins need to libmariadb-dev,
  together with libmysqlservices.a. At least until we'll
  create mariadb-plugin-dev. Nobody should need huge
  libmariadbd-dev to develop a plugin
* rpm: all in MariaDB-devel already, no changes here
* install wsrep headers too, THD layout depends on WITH_WSREP
* show DBUG_OFF, ENABLED_DEBUG_SYNC, and SAFE_MUTEX to plugins, same
  reason (it doesn't happen automatically as they're not in my_config.h)
* but don't install config.h - high chance of name conflict with other
  projects and it's an exact copy of my_config.h anyway.
* adjust plugin.cmake to work for external plugins
* move server-internal part of it to top-level CMakeLists.txt
* remove double-defined macros from unireg.h (the guard doesn't help
  if unireg.h is included first)
* package plugin metadata as yaml in .tar.gz/.zip

ColumnStore, until fixed, needs a backward-compatibility workaround
Sergei Golubchik
fix errmsg-utf8.txt dependencies for Ninja generator

GenError's custom command must specify headers as OUTPUT,
otherwise ninja cannot deduce that mysqld.cc depends on errmsg-utf8.txt

As a bonus, BYPRODUCTS lists generated files for `ninja clean`
Sergei Golubchik
fix building of external plugins

ARG_DEPENDS might be empty.

tests should be under <pluginname>/ not under whatever build dir happened
to be named.
Oleksandr Byelkin
MDEV-41094 KDF() aliases large iteration/width to weak 32-bit values

KDF() narrowed its iteration-count and key-width arguments without
checking they fit, so values differing by 2^32 aliased to the same
small value and silently derived a much weaker key.

Added checks that the key width fits a 16-bit unsigned value and the
iteration count fits a 32-bit value before narrowing them, rejecting
out-of-range values with an error instead of aliasing.
Thirunarayanan Balathandayuthapani
MDEV-19574: innodb_stats_method is not honored when innodb_stats_persistent=ON

Problem:
=======
When persistent statistics are enabled (innodb_stats_persistent=ON),
the innodb_stats_method setting is not properly utilized during
statistics calculation.

The statistics collection functions always use a hardcoded default
behavior for NULL value comparison instead of respecting the
configured stats method. This affects the accuracy of
n_diff_key_vals (distinct key count), particularly for
indexes with nullable columns containing NULL values.

Moreover, stat_n_non_null_key_vals[] was never computed for
persistent statistics; it stayed at the 0 that
dict_stats_empty_index() assigns.

With innodb_stats_method=nulls_ignored, innodb_rec_per_key()
therefore always found n_diff <= n_null and reported one record
per key for every index. This impacts the query optimizer,
which makes decisions based on inaccurate cardinality estimates.

Solution:
========
Introduced IndexLevelStats to collect statistics at a specific
B-tree level during index analysis.

Introduced PageStats to collect statistics for leaf page analysis.

Refactored the following functions:
dict_stats_analyze_index_level() to IndexLevelStats::analyze_level()
dict_stats_analyze_index_for_n_prefix() to IndexLevelStats::sample_leaf_pages()
dict_stats_analyze_index_below_cur() to PageStats::scan_below()
dict_stats_scan_page() to PageStats::scan()

The innodb_stats_method value is read once per table in
dict_stats_update_persistent() and passed down, so that all
indexes of a table are analyzed with the same method.

Add the stats method name to stat_description when
innodb_stats_method has a non-default value. The suffix is
dropped when the description is already full.

Added the new stat name n_nonnull_fld01, n_nonnull_fld02, etc.
with a stats description, to indicate how many non-null values
exist for the nth field of the index. This value is retrieved
and stored in the index statistics
in dict_stats_fetch_index_stats_step(). The counts are per
column, not per n-column prefix.

rec_get_n_blob_pages(): Calculate the number of
externally stored pages for a record, using ceiling division
by the usable BLOB page payload (blob_part_size), which differs
between ROW_FORMAT=COMPRESSED (zip_size minus FIL_PAGE_DATA) and
the other formats (srv_page_size minus the BLOB header and the
page trailer). For ROW_FORMAT=COMPRESSED the length in
the field reference is the uncompressed length, so the result is an
upper bound.

When the leaf level is scanned in full, the number of leaf pages that
were scanned is reported as n_leaf_pages for a multi level index.
Before, result.n_leaf_pages was overwritten with
index->stat_n_leaf_pages, which dict_stats_empty_index() had
just set to 1, so every index that took the full scan path
reported n_leaf_pages=1. Single page indexes report 1.
This changes cardinality estimates and
therefore leads to multiple changes in existing test cases.

Non-null values are counted only at the leaf level, since only leaf
pages hold actual records. A full scan of the leaf level counts them
exactly. When the level is sampled, the per column count is derived
from the sampled leaves with the same formula as n_diff:

  n_ordinary_leaf_pages * n_non_null_all_analyzed_pages
                        / n_leaf_pages_to_analyze

This is an estimate for NOT NULL columns as well: the sampled leaves
may hold fewer or more records than the average, and a dive that
stops at a boring page contributes nothing to the sum while still
counting in the divisor.

innodb_rec_per_key(): stat_n_non_null_key_vals[i] holds the
number of records in which the i-th indexed column alone is
not NULL, while what has to be excluded here is the number
of records whose first i+1 columns are all not NULL,
because that is the population which the n-column
prefix statistic stat_n_diff_key_vals[i] has to be
corrected against when innodb_stats_method=nulls_ignored:
with NULLs compared as unequal, every record carrying a NULL
anywhere in the prefix adds a distinct value of its own to n_diff.

PageStats::scan(): n_non_null is accumulated and assigned only
for leaf pages, so that a non-leaf scan cannot leave a node
pointer count behind when scan_below() stops at a boring page
without reaching a leaf.

IndexLevelStats::reset_for_level() also clears n_diff[], and
dict_stats_analyze_index() zero initializes the buffer backing it, so
that a level scan which finds no records (a failed
btr_pcur_open_level(), or a non-leaf page whose first record is not
marked as the leftmost one on the level) leaves n_diff[] at 0 instead
of stale values.

IndexLevelStats::sample_leaf_pages() returns early when the group
boundaries for the prefix are empty, which is the same condition.

IndexLevelStats::analyze_level(): Instead of copying the last record
of the page, retain the latch on the page until the record has been
compared with the first record of the next page

dict_stats_fetch_index_stats_step() no longer resets
stat_n_non_null_key_vals[] while processing an n_diff_pfxNN row:
dict_stats_empty_table() has already cleared the array before the
fetch, and with n_nonnull_fldNN rows now being read too,
that reset would make the result depend on the order in which the
rows arrive.

dict_stats_save(): now static function in dict0stats.cc that
takes the innodb_stats_method value, and is removed from dict0stats.h.
dict_stats_update_persistent() saves the statistics itself, so its
callers no longer have to.

Replaced btr_rec_get_externally_stored_len() with
rec_get_n_blob_pages() in dict0stats.cc.

btr_rec_get_field_ref_offs() and btr_rec_get_field_ref(),
together with the BTR_BLOB_HDR_* macros, were moved from btr0cur.cc
to btr0cur.h so that rec_get_n_blob_pages() can reuse them;

btr_rec_get_field_ref_offs() is now a noexcept function
returning size_t.

Changed stat_n_diff_key_vals and stat_n_non_null_key_vals from
ib_uint64_t* to uint64_t*

len_is_stored(): simplified to a single comparison, which is
equivalent for the unsigned lengths that it is used with.

Removed the unused UNIV_STATS_DEBUG build macro (univ.i) and
turned the DEBUG_PRINTF() helper in dict0stats.cc into an
unconditional no-op
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations (latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE) for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collation" error can be avoided in a comparison
operator. We can choose any of the sides as the operation effective
collation - the result will be equal.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPEROIRE_XXX flags set.

  This patch implements detecting tailoring equality on this reperoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st

  It returns the tailoring on the given repertoire for the given collation.

  If cs1->cset->tailoring(cs1, some_repertoire) returns {0,0},
  it means illegal mix optimization cannot be used for this collation
  on the given repertoire.

  If these calls:
    tr1= cs1->cset->tailoring(cs1, some_repertoire);
    tr2= cs2->cset->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.ptr==tr2.ptr,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Adding a new method DTCollation::aggregate_by_repertoire().

- Adding a new flag MY_COLL_ALLOW_BY_REPERTOIRE.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_REPERTOIRE.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose a collation
  of one of the sides (even if they are compatible on the given repertoire).

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of bits
  rather than a single bit, the way how to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (!(repertoire & ~MY_REPERTOIRE_ASCII))

- New flags were added int for CHARSET_INFO::state
  * MY_CS_ASCII_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation does not has no irregularities
    on the ASCII range.

  * MY_CS_IDENT_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has not irregularities
    on the IDENT subrange only (but can have irregularities say on
    punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII letters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_BINARY_CI and MY_CS_IDENT_BINARY_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as returned by:
    cs->cset->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests.

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.

- Adding a number of MTR tests in the form of the originally
  reported stript for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';
Alessandro Vetere
MDEV-39792 InnoDB: ALTER TABLE FORCE triggers assertion "s" in buf_page_get_gen()

When rebuilding a table from ROW_FORMAT=COMPACT or DYNAMIC into
ROW_FORMAT=REDUNDANT, row_merge_buf_add() fetches the full value of an
externally stored (off-page) CHAR column in a multi-byte character set
and pads it to REDUNDANT's fixed local width via
row_merge_buf_redundant_convert(). That helper already dereferences the
BLOB and calls dfield_set_data(), which clears the field's "externally
stored" flag, since the value is now held in full locally.

The "flag externally stored fields" step further down in
row_merge_buf_add() did not know this had happened. It still consulted
the row_ext_t cache built from the original (pre-conversion) record and,
for a column that is not part of the clustered index's unique key,
called dfield_set_ext() again on the very field that had just been
converted, without restoring its data pointer to a valid 20-byte
external reference. row_merge_copy_blobs() would then read the tail of
the padded, space-filled buffer as if it were a BTR_EXTERN_FIELD_REF,
deriving a garbage tablespace id and crashing buf_page_get_gen()'s
fil_space_get() assertion when the alter tried to build the new
clustered index.

Skip the re-flagging step for a field whose "externally stored" flag is
no longer set. row_build() flags every off-page column, and the
row_ext_t cache only holds a subset of those columns, so a field that
is not flagged is either a converted one (already fully local) or one
that the cache does not hold.

With the field no longer re-flagged, the rebuild completes, and the
rebuilt table passes CHECK TABLE with the full column value.

The MDEV-31025 case in innodb.default_row_format_alter failed on
innodb_page_size=4k and 8k: its ROW_FORMAT=REDUNDANT table has eight
utf32 CHAR(255) columns, which CREATE TABLE rejects with
ER_TOO_BIG_ROWSIZE on those page sizes. Derive the number of columns
from the page size, so that the record still exceeds the maximum local
record size and the fixed-length column c is stored externally. The
whole test now passes on every page size.
ParadoxV5
MDEV-38849 slave_connections_needed_for_purge prevents independent machine from purging binary logs

`@@slave_connections_needed_for_purge`’s default of `1` ensures binary
log availability on replication masters, but is not a sensible default
suitable for all scenarios, especially for long-term slave servers and
standalone (not in a replication setup) servers.
The outcome was that standalone server users were confused why automatic
binlog purging does not work.

This commit changes this default to `0`, which is suitable for both
standalone and (when backed by prompt failure recovery)
replication setups.
`0` also more closely matches the behaviour before MDEV-31404,
which added this variable, out of the box.

This commit also adds a one-time replication warning when registering a
slave, but `@@slave_connections_needed_for_purge` is left unchanged.
Rather than enforcing a defence with an unsensible default, this
reminder will bring awareness of the risk of automatic binlog purging.

This commit also cleans up Galera and MTR workarounds to the
introduction of the `@@slave_connections_needed_for_purge=1` default.