Home - Waterfall Grid T-Grid Console Builders Recent Builds Buildslaves Changesources - JSON API - About

Console View


Categories: connectors experimental galera main
Legend:   Passed Failed Warnings Failed Again Running Exception Offline No data

connectors experimental galera main
Yuchen Pei
MDEV-41366 Check prefix key match in iterated partition index scan

The idea of Case 2 in can_skip_merging_scans is that when partition
column key prefix is fixed, and the infix that is also the "partition
by range" column, we can scan each partition in order. This relies on
accurately returning EOF when a partition has no (more) matching rows.
It is possible to have no more matching rows at the first index read
of a partition in a non-exact access, in which case we need to check
the prefix match to rule out false positives and return EOF correctly.

This change by itself would result in incomplete results, if Case 2
incorrectly determines the scan can iterate over partitions in order
in AFTER_KEY and BEFORE_KEY reads. These reads when applied to a
partition column key prefix, necessarily means seeking a different
prefix value. This could result in false negatives. To that end, we
strengthen the checks in Case 2 so that AFTER_KEY and BEFORE_KEY reads
require both the prefix and the partitioned column itself in the
keypart_map to ensure correctness.

For example, this check disqualifies the index_read_map call in an
existing loose index scan testcase, causing the index scan to use both
unordered (in read_range_first) and ordered.

Since we have to enumerate all flags, we include a fix of missing
reverse_order for KEY_OR_PREV. The only known usecase is HANDLER ...
READ which is not working correctly and due to be fixed in MDEV-41364.
For consistency with other equally incorrect HANDLER cases involving
other flags, we do not single out KEY_OR_PREV for special
treatment (e.g. return false in can_skip_merging_scans).
Yuchen Pei
MDEV-41366 Check prefix key match in iterated partition index scan

The idea of Case 2 in can_skip_merging_scans is that when partition
column key prefix is fixed, and the infix that is also the "partition
by range" column, we can scan each partition in order. This relies on
accurately returning EOF when a partition has no (more) matching rows.
It is possible to have no more matching rows at the first index read
of a partition in a non-exct access, in which case we need to check
the prefix match to rule out false positives and return EOF correctly.

This change by itself would result in incomplete results, if Case 2
incorrectly determines the scan can iterate over partitions in order
in AFTER_KEY and BEFORE_KEY reads. These reads when applied to a
partition column key prefix, necessarily means seeking a different
prefix value. This could result in false negatives. To that end, we
strengthen the checks in Case 2 so that AFTER_KEY and BEFORE_KEY reads
require both the prefix and the partitioned column itself in the
keypart_map to ensure correctness.

For example, this check disqualifies the index_read_map call in an
existing loose index scan testcase, causing the index scan to use both
unordered (in read_range_first) and ordered.
Yuchen Pei
MDEV-41366 Check prefix key match in iterated partition index scan

The idea of Case 2 in can_skip_merging_scans is that when partition
column key prefix is fixed, and the infix that is also the "partition
by range" column, we can scan each partition in order. This relies on
accurately returning EOF when a partition has no (more) matching rows.
It is possible to have no more matching rows at the first index read
of a partition in a non-exact access, in which case we need to check
the prefix match to rule out false positives and return EOF correctly.

This change by itself would result in incomplete results, if Case 2
incorrectly determines the scan can iterate over partitions in order
in AFTER_KEY and BEFORE_KEY reads. These reads when applied to a
partition column key prefix, necessarily means seeking a different
prefix value. This could result in false negatives. To that end, we
strengthen the checks in Case 2 so that AFTER_KEY and BEFORE_KEY reads
require both the prefix and the partitioned column itself in the
keypart_map to ensure correctness.

For example, this check disqualifies the index_read_map call in an
existing loose index scan testcase, causing the index scan to use both
unordered (in read_range_first) and ordered.

Since we have to enumerate all flags, we include a fix of missing
reverse_order for KEY_OR_PREV. The only known usecase is HANDLER ...
READ which is not working correctly and due to be fixed in MDEV-41364.
For consistency with other equally incorrect HANDLER cases involving
other flags, we do not single out KEY_OR_PREV for special
treatment (e.g. return false in can_skip_merging_scans).
Sergei Golubchik
cleanup: extract duplicated no_part_keypart() check
Marko Mäkelä
MDEV-41395 fil_space_t::create() may release fil_system.mutex

If fil_space_t::create() ends up releasing fil_system.mutex for
invoking fil_crypt_threads_signal(), another thread may grab
fil_system.mutex and find a tablespace that contains no data files.
Let us fix this race condition and remove some checks that now
are redundant.

fil_space_t::create(): Do not invoke fil_crypt_threads_init();
let the caller do that. During database startup, before
fil_crypt_threads_init() has been invoked, there is no point to
invoke this function; the encryption queues will be properly
set up for the tablespaces that have been opened by that time.

fil_ibd_create(): Invoke fil_crypt_threads_signal() after
releasing fil_system.mutex.

fil_crypt_threads_signal(): Check that fil_crypt_threads_init()
has been called. The function dict_table_open_on_name() will be
invoked during early server startup.

dict_table_open_on_name(): Invoke fil_crypt_threads_init()
at the end if the table was loaded into the cache and the
tablespace likely as well. This should be the main code path
through which old tablespaces are opened via SQL.

We will not invoke fil_crypt_threads_init() it in
dict_load_table_on_id() or dict_sys.load_table(); it will instead
be invoked at a higher level, typically when a table definition is
first loaded by name, in dict_table_open_on_name().

buf_dump_load_func(): After the non-first buf_load(), invoke
fil_crypt_threads_signal() in order to start the encryption
of any newly loaded tablespaces.

drop_garbage_tables_after_restore(), wsrep_append_foreign_key(),
ha_innobase::create(), ha_innobase::truncate(),
row_rename_table_for_mysql(), i_s_sys_tables_fill_table(),
i_s_sys_tables_fill_table_stats(): Invoke fil_crypt_threads_signal()
in case some tablespace metadata was loaded.

row_import_cleanup(): Invoke fil_crypt_threads_signal() to register
the imported tablespace.

i_s_sys_tablespaces_fill_table(), fil_system_t::find(),
Datafile::validate_first_page(), fil_space_t::name(),
fil_space_t::read_page0(), fil_space_t::prepare_acquired(),
fil_space_t::try_to_close(), fil_crypt_default_encrypt_tables_fill(),
fil_system_t::default_encrypt_next(): Remove a now-redundant check
for a tablespace that contains no files.
Brad Smith
MDEV-41221: memcpy() source and target overlap in dict_table_rename_in_cache()

In dict_table_rename_in_cache(), sql_id = foreign->sql_id() points into
the existing foreign->id string. When the new id is not longer than the old
one (always the case when renaming from "#sql-alter-..." back to the real
table name), the buffer is reused in place, so the final
snprintf(id, fklen, "%s\377%s", table->name.m_name, sql_id)
reads sql_id from the buffer it is writing to (sql_id == id + 50 in gdb).
That is undefined behavior. Most libcs happen to copy forward and get away
with it, but OpenBSD's memcpy detects the overlap and calls abort().

The fix computes the lengths explicitly, memmove()s the sql_id tail to its
final position first, then memcpy()s the table name prefix and writes the
\377 separator. The decision to reuse the old buffer or allocate a new one
is unchanged.

Co-authored-by: Claude Opus 5.5 <[email protected]>
Yuchen Pei
MDEV-41322 Convert KEY_NOT_FOUND to EOF in "unordered" partition index scans during index_next[_same]/index_prev calls

ha_partition::handle_unordered_scan_next_partition is called in a
variety of accesses, including index_read, index_prev, and index_next.

ha_partition::handle_unordered_next and
ha_partition::handle_unordered_prev are called from index_next and
index_prev accesses. They check for signs of end of
scan (m_part_spec.start_part == NO_CURRENT_PART_ID) and exit early if
that is the case. HA_ERR_END_OF_FILE (EOF) is commonly paired with
assigning NO_CURRENT_PART_ID to m_part_spec.start_part whenever
possible, to signal the end of scan.

The error HA_ERR_KEY_NOT_FOUND means the requested key is not found.
It should not mean the end of scan, when for example
ha_partition::handle_unordered_scan_next_partition is called from
index_read, because a subsequent index_next[_same] / index_prev call
would then incorrectly return immediately from
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev. This is why HA_ERR_KEY_NOT_FOUND
is retained and returned in
ha_partition::handle_unordered_scan_next_partition.

But if the call is from index_next[_same] / index_prev,
HA_ERR_KEY_NOT_FOUND should indeed mean end of scan. In this patch, we
ensure this is the case by converting HA_ERR_KEY_NOT_FOUND to EOF in
ha_partition::handle_unordered_next /
ha_partition::handle_unordered_prev.
Sergei Golubchik
MDEV-25848 Support for Multi-Valued Indexes

* New type of high-level indexes - JSON ARRAY
* sql/index/json.{h,cc}
  * table structure (value varbinary(255), tref varbinary(N), PK(value,tref))
  * insert/update/delete/search works, search uses keyread
  * on insert json array is parsed, values are inserted
    * prefixed by the type tag, only simple types are supported
  * records_in_range works, but only for equalities (MEMBER OF)
* MySQL-compatible syntax `INDEX ((CAST(col AS type ARRAY)))`, but
  * col is a bare name, not an expression, expressions are part of MDEV-35853
  * type is ignored, not stored, also part of MDEV-35853
  * thus `'[1, 0.2, true, "foo", null]'` works fine, type not enforced
  * thus SHOW CREATE TABLE prints `... AS JSON ARRAY`
* `const MEMBER OF (col)` uses index, via
  * range optimizer, QUICK_RANGE_SELECT_ARRAY
  * Item_func_member_of::get_mm_tree()/etc
Yuchen Pei
MDEV-41366 Check prefix key match in iterated partition index scan

The idea of Case 2 in can_skip_merging_scans is that when partition
column key prefix is fixed, and the infix that is also the "partition
by range" column, we can scan each partition in order. This relies on
accurately returning EOF when a partition has no (more) matching rows.
It is possible to have no more matching rows at the first index read
of a partition in a non-exact access, in which case we need to check
the prefix match to rule out false positives and return EOF correctly.

This change by itself would result in incomplete results, if Case 2
incorrectly determines the scan can iterate over partitions in order
in AFTER_KEY and BEFORE_KEY reads. These reads when applied to a
partition column key prefix, necessarily means seeking a different
prefix value. This could result in false negatives. To that end, we
strengthen the checks in Case 2 so that AFTER_KEY and BEFORE_KEY reads
require both the prefix and the partitioned column itself in the
keypart_map to ensure correctness.

For example, this check disqualifies the index_read_map call in an
existing loose index scan testcase, causing the index scan to use both
unordered (in read_range_first) and ordered.

Since we have to enumerate all flags, we include a fix of missing
reverse_order for KEY_OR_PREV. The only known usecase is HANDLER ...
READ which is not working correctly and due to be fixed in MDEV-41364.
For consistency with other equally incorrect HANDLER cases involving
other flags, we do not single out KEY_OR_PREV for special
treatment (e.g. return false in can_skip_merging_scans).
Marko Mäkelä
MDEV-41395 fil_space_t::create() may release fil_system.mutex

If fil_space_t::create() ends up releasing fil_system.mutex for
invoking fil_crypt_threads_signal(), another thread may grab
fil_system.mutex and find a tablespace that contains no data files.
Let us fix this race condition and remove some checks that now
are redundant.

fil_space_t::create(): Do not invoke fil_crypt_threads_init();
let the caller do that. During database startup, before
fil_crypt_threads_init() has been invoked, there is no point to
invoke this function; the encryption queues will be properly
set up for the tablespaces that have been opened by that time.

fil_ibd_create(): Invoke fil_crypt_threads_signal() after
releasing fil_system.mutex.

fil_crypt_threads_signal(): Check that fil_crypt_threads_init()
has been called. The function dict_table_open_on_name() will be
invoked during early server startup.

dict_table_open_on_name(): Invoke fil_crypt_threads_init()
at the end if the table was loaded into the cache and the
tablespace likely as well. This should be the main code path
through which old tablespaces are opened via SQL.

We will not invoke fil_crypt_threads_init() in
dict_load_table_on_id() or dict_sys.load_table(); it will instead
be invoked at a higher level, typically when a table definition is
first loaded by name, in dict_table_open_on_name().

buf_dump_load_func(): After the non-first buf_load(), invoke
fil_crypt_threads_signal() in order to start the encryption
of any newly loaded tablespaces.

drop_garbage_tables_after_restore(), wsrep_append_foreign_key(),
ha_innobase::create(), ha_innobase::truncate(),
row_rename_table_for_mysql(), i_s_sys_tables_fill_table(),
i_s_sys_tables_fill_table_stats(): Invoke fil_crypt_threads_signal()
in case some tablespace metadata was loaded.

row_import_cleanup(): Invoke fil_crypt_threads_signal() to register
the imported tablespace.

i_s_sys_tablespaces_fill_table(), fil_system_t::find(),
Datafile::validate_first_page(), fil_space_t::name(),
fil_space_t::read_page0(), fil_space_t::prepare_acquired(),
fil_space_t::try_to_close(), fil_crypt_default_encrypt_tables_fill(),
fil_system_t::default_encrypt_next(): Remove a now-redundant check
for a tablespace that contains no files.

(cherry picked from commit 8214825d675167a13f9dabaebf11674d21a26339)
Marko Mäkelä
MDEV-41395 fil_space_t::create() may release fil_system.mutex

If fil_space_t::create() ends up releasing fil_system.mutex for
invoking fil_crypt_threads_signal(), another thread may grab
fil_system.mutex and find a tablespace that contains no data files.
Let us fix this race condition and remove some checks that now
are redundant.

fil_space_t::create(): Do not invoke fil_crypt_threads_init();
let the caller do that. During database startup, before
fil_crypt_threads_init() has been invoked, there is no point to
invoke this function; the encryption queues will be properly
set up for the tablespaces that have been opened by that time.

fil_ibd_create)(: Invoke fil_crypt_threads_signal() after
releasing fil_system.mutex.

fil_crypt_threads_signal(): Check that fil_crypt_threads_init()
has been called. The function dict_table_open_on_name() will be
invoked during early server startup.

dict_table_open_on_name(): Invoke fil_crypt_threads_init()
at the end if the table was loaded into the cache and the
tablespace likely as well. This should be the main code path
through which old tablespaces are opened via SQL.

We will not invoke fil_crypt_threads_init() it in
dict_load_table_on_id() or dict_sys.load_table(); it will instead
be invoked at a higher level, typically when a table definition is
first loaded by name, in dict_table_open_on_name().

buf_dump_load_func(): After the non-first buf_load(), invoke
fil_crypt_threads_signal() in order to start the encryption
of any newly loaded tablespaces.

drop_garbage_tables_after_restore(), wsrep_append_foreign_key(),
ha_innobase::create(), ha_innobase::truncate(),
row_rename_table_for_mysql(), i_s_sys_tables_fill_table(),
i_s_sys_tables_fill_table_stats(): Invoke fil_crypt_threads_signal()
in case some tablespace metadata was loaded.

row_import_cleanup(): Invoke fil_crypt_threads_signal() to register
the imported tablespace.

i_s_sys_tablespaces_fill_table(), fil_system_t::find(),
Datafile::validate_first_page(), fil_space_t::name(),
fil_space_t::read_page0(), fil_space_t::prepare_acquired(),
fil_space_t::try_to_close(), fil_crypt_default_encrypt_tables_fill(),
fil_system_t::default_encrypt_next(): Remove a now-redundant check
for a tablespace that contains no files.
Sergei Golubchik
cleanup: remove redundant function
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations
(latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE)
for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collations" error can be avoided in
a comparison operator. We can choose any of the sides as the operation
effective collation - the result will be equal.

This optimization is not applied when at least one side has an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison are
already illegal when the character sets are the same, so for
consistency this stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality
  on these repertoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st.

  It returns the tailoring on the given repertoire for the given
  collation.

  If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0},
  it means the illegal mix optimization cannot be used
  for this collation on the given repertoire.

  If these calls:
    tr1= cs1->coll->tailoring(cs1, some_repertoire);
    tr2= cs2->coll->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.str==tr2.str,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Tailoring strings are shared constants my_tailoring_str_*
  (strings/strings_def.h). Tailorings are compared by the string
  pointer, so tailoring() implementations return strings only from
  these constants. The strings do not depend on PAD/NOPAD.

- Adding a new method DTCollation::aggregate_by_tailoring().
  Collations with different MY_CS_NOPAD are never compatible.
  UCA collations of different versions are not compatible on
  the repertoire with ASCII punctuation, as the UCA version
  affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT
  and CIRCUMFLEX ACCENT after PERCENT SIGN).
  Like the other successful branches of DTCollation::aggregate(),
  it makes the resulting repertoire cover both sides.

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose
  a collation of one of the sides (even if they are compatible
  on the given repertoire). This also applies to ExtractValue()
  and UpdateXML(), which now aggregate their arguments without
  this flag.

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of
  bits rather than a single bit, the way to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII))

  The new function my_repertoire_is_subset_of() in m_ctype.h
  is used for this purpose in C code. In C++ code
  DTCollation::repertoire_is_subset_of() is used, e.g.:

    if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII))

- Repertoire of some expressions was adjusted to the new meaning:
  * Item_null now has MY_REPERTOIRE_NONE (was ASCII).
  * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30
    (was EXTENDED).
  * Lex_string_with_metadata_st::repertoire(cs) now scans the string
    contents to detect the actual repertoire.
  * HEX() now has MY_REPERTOIRE_ASCII_ALNUM.
  * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY,
    JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra
    characters they put into the result (quotes, separators, padding,
    brackets).
  * LOWER() and UPPER() add the letters of both cases to the repertoire,
    because the result can contain letters of the opposite case.
    In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an
    ASCII letter can be converted to a non-ASCII one (I -> dotless i).
    Unicode collations are detected by the new member
    casefold_info_st::can_convert_to_non_ascii_on_casefolding,
    simple 8bit collations by to_lower['I'] and to_upper['i'].

- The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL),
  so my_charset_is_ascii_based() is true for tis620.
  This changes the result of subselect_extra_no_semijoin
  from an error to success.

- tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case,
  so the underscore sorts before the letters, unlike in the IDENT
  tailoring. These collations now exclude the IDENT repertoire:
  the tailoring optimization is only allowed for them on ALNUM.

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII
    repertoire, and converts lower case letters to upper case ones,
    so the underscore sorts after the letters (with an upper to lower
    case conversion it would sort before the letters).

  * MY_CS_IDENT_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT
    repertoire (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII characters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in
  plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as
  returned by:
    cs->coll->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests. Tailorings do not depend on PAD/NOPAD
  and on the UCA version, so the function shows "[nopad]" for NOPAD
  collations, and "[version X.Y.Z]" for UCA collations on the ASCII
  repertoire. The version is formatted by a new class UCAVersion,
  which uses a new method CharBuffer::append_uint8().

- Adding the test plugin/func_test/mysql-test/func_test/
  ctype_caseup_ci_verify.test. It checks the order of ASCII characters
  for all collations declared "caseup CI" on the ASCII or IDENT range.

- Adding the test plugin/func_test/mysql-test/func_test/
  collation_tailoring_args.test. It checks the arguments of the plugin
  function collation_tailoring(): the repertoire passed by name or by
  number, and the NULL results for unknown names.

- Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test.
  It displays the tailorings of collations loaded at runtime from
  Index.xml: simple 8bit collations (flags detected from the weights)
  and Unicode collations (LDML rules which do or do not reorder ASCII).

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';

  The repeated blocks of queries are shared through
  mysql-test/include/
  ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc.

- Adding tests to ctype_cp932.test checking that ExtractValue() and
  UpdateXML() still raise "Illegal mix of collations".
ParadoxV5
MDEV-40271: Disallow `\0` and `\n` in replication connection names

Null bytes conflict with strings’ null-termination,
\while newline bytes are delimiters in the `multi-master.info` file.
While filename-unsafe chars are canonicalized when mapping names
to filenames, these two still cause trouble in other places.
Therefore, this commit extends `check_master_connection_name()`
to also reject names containing these chars.

This check function is now also used for system variable basenames
(e.g., in `@@replicate_do_db`), which (probably harmlessly)
did not have the existing `MAX_CONNECTION_NAME` check.
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations
(latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE)
for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collations" error can be avoided in
a comparison operator. We can choose any of the sides as the operation
effective collation - the result will be equal.

This optimization is not applied when at least one side has an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison are
already illegal when the character sets are the same, so for
consistency this stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality
  on these repertoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st.

  It returns the tailoring on the given repertoire for the given
  collation.

  If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0},
  it means the illegal mix optimization cannot be used
  for this collation on the given repertoire.

  If these calls:
    tr1= cs1->coll->tailoring(cs1, some_repertoire);
    tr2= cs2->coll->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.str==tr2.str,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Tailoring strings are shared constants my_tailoring_str_*
  (strings/strings_def.h). Tailorings are compared by the string
  pointer, so tailoring() implementations return strings only from
  these constants. The strings do not depend on PAD/NOPAD.

- Adding a new method DTCollation::aggregate_by_tailoring().
  Collations with different MY_CS_NOPAD are never compatible.
  UCA collations of different versions are not compatible on
  the repertoire with ASCII punctuation, as the UCA version
  affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT
  and CIRCUMFLEX ACCENT after PERCENT SIGN).
  Like the other successful branches of DTCollation::aggregate(),
  it makes the resulting repertoire cover both sides.

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose
  a collation of one of the sides (even if they are compatible
  on the given repertoire). This also applies to ExtractValue()
  and UpdateXML(), which now aggregate their arguments without
  this flag.

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of
  bits rather than a single bit, the way to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII))

  The new function my_repertoire_is_subset_of() in m_ctype.h
  is used for this purpose in C code. In C++ code
  DTCollation::repertoire_is_subset_of() is used, e.g.:

    if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII))

- Repertoire of some expressions was adjusted to the new meaning:
  * Item_null now has MY_REPERTOIRE_NONE (was ASCII).
  * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30
    (was EXTENDED).
  * Lex_string_with_metadata_st::repertoire(cs) now scans the string
    contents to detect the actual repertoire.
  * HEX() now has MY_REPERTOIRE_ASCII_ALNUM.
  * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY,
    JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra
    characters they put into the result (quotes, separators, padding,
    brackets).
  * LOWER() and UPPER() add the letters of both cases to the repertoire,
    because the result can contain letters of the opposite case.
    In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an
    ASCII letter can be converted to a non-ASCII one (I -> dotless i).
    Unicode collations are detected by the new member
    casefold_info_st::can_convert_to_non_ascii_on_casefolding,
    simple 8bit collations by to_lower['I'] and to_upper['i'].

- The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL),
  so my_charset_is_ascii_based() is true for tis620.
  This changes the result of subselect_extra_no_semijoin
  from an error to success.

- tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case,
  so the underscore sorts before the letters, unlike in the IDENT
  tailoring. These collations now exclude the IDENT repertoire:
  the tailoring optimization is only allowed for them on ALNUM.

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII
    repertoire.

  * MY_CS_IDENT_BINARY_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT
    repertoire (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII characters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_BINARY_CI and MY_CS_IDENT_BINARY_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in
  plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as
  returned by:
    cs->coll->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests. Tailorings do not depend on PAD/NOPAD
  and on the UCA version, so the function shows "[nopad]" for NOPAD
  collations, and "[version X.Y.Z]" for UCA collations on the ASCII
  repertoire. The version is formatted by a new class UCAVersion,
  which uses a new method CharBuffer::append_uint8().

- Adding the test plugin/func_test/mysql-test/func_test/
  ctype_binary_ci_verify.test. It checks the order of ASCII characters
  for all collations declared "binary CI" on the ASCII or IDENT range.

- Adding the test plugin/func_test/mysql-test/func_test/
  collation_tailoring_args.test. It checks the arguments of the plugin
  function collation_tailoring(): the repertoire passed by name or by
  number, and the NULL results for unknown names.

- Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test.
  It displays the tailorings of collations loaded at runtime from
  Index.xml: simple 8bit collations (flags detected from the weights)
  and Unicode collations (LDML rules which do or do not reorder ASCII).

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';

  The repeated blocks of queries are shared through
  mysql-test/include/
  ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc.

- Adding tests to ctype_cp932.test checking that ExtractValue() and
  UpdateXML() still raise "Illegal mix of collations".
drrtuy
MDEV-40957 MTR  tests to cover DML statements in a replicated setup.
Sergei Golubchik
Merge branch '11.4' into 11.8
Sergei Golubchik
cleanup: is_local_field() -> get_local_field()

removes the need for a separate cast after the check,
and guarantees that the correct item is used as Item_field
(and not something else was mistakenly cast due to a typo)
Marko Mäkelä
MDEV-41395 fil_space_t::create() may release fil_system.mutex

If fil_space_t::create() ends up releasing fil_system.mutex for
invoking fil_crypt_threads_signal(), another thread may grab
fil_system.mutex and find a tablespace that contains no data files.
Let us fix this race condition and remove some checks that now
are redundant.

fil_space_t::create(): Do not invoke fil_crypt_threads_init();
let the caller do that. During database startup, before
fil_crypt_threads_init() has been invoked, there is no point to
invoke this function; the encryption queues will be properly
set up for the tablespaces that have been opened by that time.

fil_ibd_create(): Invoke fil_crypt_threads_signal() after
releasing fil_system.mutex.

fil_crypt_threads_signal(): Check that fil_crypt_threads_init()
has been called. The function dict_table_open_on_name() will be
invoked during early server startup.

dict_table_open_on_name(): Invoke fil_crypt_threads_init()
at the end if the table was loaded into the cache and the
tablespace likely as well. This should be the main code path
through which old tablespaces are opened via SQL.

We will not invoke fil_crypt_threads_init() in
dict_load_table_on_id() or dict_sys.load_table(); it will instead
be invoked at a higher level, typically when a table definition is
first loaded by name, in dict_table_open_on_name().

buf_dump_load_func(): After the non-first buf_load(), invoke
fil_crypt_threads_signal() in order to start the encryption
of any newly loaded tablespaces.

drop_garbage_tables_after_restore(), wsrep_append_foreign_key(),
ha_innobase::create(), ha_innobase::truncate(),
row_rename_table_for_mysql(), i_s_sys_tables_fill_table(),
i_s_sys_tables_fill_table_stats(): Invoke fil_crypt_threads_signal()
in case some tablespace metadata was loaded.

row_import_cleanup(): Invoke fil_crypt_threads_signal() to register
the imported tablespace.

i_s_sys_tablespaces_fill_table(), fil_system_t::find(),
Datafile::validate_first_page(), fil_space_t::name(),
fil_space_t::read_page0(), fil_space_t::prepare_acquired(),
fil_space_t::try_to_close(), fil_crypt_default_encrypt_tables_fill(),
fil_system_t::default_encrypt_next(): Remove a now-redundant check
for a tablespace that contains no files.
Sergei Golubchik
cleanup: remove obsolete code
Yuchen Pei
MDEV-41366 Check prefix key match in iterated partition index scan

The idea of Case 2 in can_skip_merging_scans is that when partition
column key prefix is fixed, and the infix that is also the "partition
by range" column, we can scan each partition in order. This relies on
accurately returning EOF when a partition has no (more) matching rows.
It is possible to have no more matching rows at the first index read
of a partition in a non-exact access, in which case we need to check
the prefix match to rule out false positives and return EOF correctly.

This change by itself would result in incomplete results, if Case 2
incorrectly determines the scan can iterate over partitions in order
in AFTER_KEY and BEFORE_KEY reads. These reads when applied to a
partition column key prefix, necessarily means seeking a different
prefix value. This could result in false negatives. To that end, we
strengthen the checks in Case 2 so that AFTER_KEY and BEFORE_KEY reads
require both the prefix and the partitioned column itself in the
keypart_map to ensure correctness.

For example, this check disqualifies the index_read_map call in an
existing loose index scan testcase, causing the index scan to use both
unordered (in read_range_first) and ordered.

Since we have to enumerate all flags, we include a fix of missing
reverse_order for KEY_OR_PREV. The only known usecase is HANDLER ...
READ which is not working correctly and due to be fixed in MDEV-41364.
For consistency with other equally incorrect HANDLER cases involving
other flags, we do not single out KEY_OR_PREV for special
treatment (e.g. return false in can_skip_merging_scans).
Sergei Golubchik
cleanup: remove dead code
forkfun
MDEV-31527 Flush buffered early messages on early exit

Buffered_logs::flush() prints and frees the buffered early-option
messages. Call it on every early exit, so warnings and errors are
not lost and the buffer is not leaked.

Make the early-option error exit in mysqld_main() unconditional.
Without the performance schema, a failed early option parse left
"----file-marker----" in argv, and it was reported as an unknown
option.

Index.xml parse errors go through error_handler_hook, not
vprint_msg_to_log(), so --validate-config reported
"Configuration is valid." Install a hook that sets
validate_config_has_warnings.

Test cleanup.
Alexander Barkov
MDEV-41246 "Illegal mix of collations" on the mysql.user view

This SQL script failed:

SET NAMES latin1 COLLATE latin1_swedish_ci;
CREATE OR REPLACE VIEW v1 AS SELECT 'Y' AS c1;
SET NAMES big5 COLLATE big5_chinese_ci;
SELECT * FROM v1 WHERE c1='y';

with the following error:

ERROR 1267 (HY000): Illegal mix of collations
(latin1_swedish_ci,COERCIBLE) and (big5_chinese_ci,COERCIBLE)
for operation '='

Note, latin1_swedish_ci and big5_chinese_ci are used here as examples.
The error also happened with different collation combinations.

Fix main idea:

If two collations have equal comparison rules (known as "tailoring")
on a given character repertoire,
like latin1_swedish_ci and big5_chinese_ci on ASCII letters,
then the "Illegal mix of collations" error can be avoided in
a comparison operator. We can choose any of the sides as the operation
effective collation - the result will be equal.

This optimization is not applied when at least one side has an explicit
COLLATE clause. Two explicit COLLATE clauses in one comparison are
already illegal when the character sets are the same, so for
consistency this stays illegal when the character sets differ too.

Most important details:

- Splitting enum_repertoire_t into smaller subsets,
  for better repertoire granularity.
  A variable holding a repertoire value can now have multiple
  MY_REPERTOIRE_XXX flags set.

  This patch implements detecting tailoring equality
  on these repertoires:
  * MY_REPERTOIRE_ASCII_ALNUM - [A..Z,a..z,0..9].
  * MY_REPERTOIRE_ASCII_IDENT - ALNUM + underscore
  * MY_REPERTOIRE_ASCII      - the entire range U+0000..U+007F

- Adding a new virtual function "tailoring" in my_collation_handler_st.

  It returns the tailoring on the given repertoire for the given
  collation.

  If cs1->coll->tailoring(cs1, some_repertoire) returns {0,0},
  it means the illegal mix optimization cannot be used
  for this collation on the given repertoire.

  If these calls:
    tr1= cs1->coll->tailoring(cs1, some_repertoire);
    tr2= cs2->coll->tailoring(cs2, some_repertoire);
  return both non-NULL results and tr1.str==tr2.str,
  then these collations are equal on the given repertoire
  and are mutually replaceable for a comparison operator,
  so "Illegal mix of collations" can be avoided.

- Tailoring strings are shared constants my_tailoring_str_*
  (strings/strings_def.h). Tailorings are compared by the string
  pointer, so tailoring() implementations return strings only from
  these constants. The strings do not depend on PAD/NOPAD.

- Adding a new method DTCollation::aggregate_by_tailoring().
  Collations with different MY_CS_NOPAD are never compatible.
  UCA collations of different versions are not compatible on
  the repertoire with ASCII punctuation, as the UCA version
  affects its order (e.g. UCA-6.2.0 moved GRAVE ACCENT
  and CIRCUMFLEX ACCENT after PERCENT SIGN).
  Like the other successful branches of DTCollation::aggregate(),
  it makes the resulting repertoire cover both sides.

- Adding a new flag MY_COLL_ALLOW_BY_TAILORING.
  It indicates to DTCollation::aggregate() that the illegal
  mix optimization by repertoire can be used in the given context.

  MY_COLL_CMP_CONV now includes MY_COLL_ALLOW_BY_TAILORING.
  Note, only comparison operators pass this flag.
  Functions returning a string result do not pass this flag,
  because in operations like CONCAT(a,b) we still need to evaluate
  precisely the collation of the result - we cannot just choose
  a collation of one of the sides (even if they are compatible
  on the given repertoire). This also applies to ExtractValue()
  and UpdateXML(), which now aggregate their arguments without
  this flag.

- As in my_repertoire_t the value MY_REPERTOIRE_ASCII is now a set of
  bits rather than a single bit, the way to detect "is only ASCII"
  repertoires has changed in the code.

  For example:
    // repertoire *IS* ascii
    if (repertoire == MY_REPERTOIRE_ASCII)

  has changed in multiple places in the code to

    // repertoire *HAS* only ascii characters
    if (my_repertoire_is_subset_of(repertoire, MY_REPERTOIRE_ASCII))

  The new function my_repertoire_is_subset_of() in m_ctype.h
  is used for this purpose in C code. In C++ code
  DTCollation::repertoire_is_subset_of() is used, e.g.:

    if (collation.repertoire_is_subset_of(MY_REPERTOIRE_ASCII))

- Repertoire of some expressions was adjusted to the new meaning:
  * Item_null now has MY_REPERTOIRE_NONE (was ASCII).
  * MY_LOCALE::repertoire() now returns MY_REPERTOIRE_UNICODE30
    (was EXTENDED).
  * Lex_string_with_metadata_st::repertoire(cs) now scans the string
    contents to detect the actual repertoire.
  * HEX() now has MY_REPERTOIRE_ASCII_ALNUM.
  * DATE_FORMAT() decides on MY_REPERTOIRE_EXTENDED using the locale
    which is actually used: the explicit third argument, or
    @@lc_time_names. It was always @@lc_time_names before.
    A non-constant locale argument is assumed to be non-ASCII.
  * QUOTE, MAKE_SET, EXPORT_SET, LPAD, RPAD, GROUP_CONCAT, JSON_ARRAY,
    JSON_OBJECT and JSON_OBJECTAGG add the repertoire of the extra
    characters they put into the result (quotes, separators, padding,
    brackets).
  * LOWER() and UPPER() add the letters of both cases to the repertoire,
    because the result can contain letters of the opposite case.
    In Turkish collations they also add MY_REPERTOIRE_EXTENDED, as an
    ASCII letter can be converted to a non-ASCII one (I -> dotless i).
    Unicode collations are detected by the new member
    casefold_info_st::can_convert_to_non_ascii_on_casefolding,
    simple 8bit collations by to_lower['I'] and to_upper['i'].

- The tis620 collations now set CHARSET_INFO::tab_to_uni (was NULL),
  so my_charset_is_ascii_based() is true for tis620.
  This changes the result of subselect_extra_no_semijoin
  from an error to success.

- tis620_thai_ci and tis620_thai_nopad_ci fold letters to lower case,
  so the underscore sorts before the letters, unlike in the IDENT
  tailoring. These collations now exclude the IDENT repertoire:
  the tailoring optimization is only allowed for them on ALNUM.

- New flags were added for CHARSET_INFO::state
  * MY_CS_ASCII_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the ASCII
    repertoire, and converts lower case letters to upper case ones,
    so the underscore sorts after the letters (with an upper to lower
    case conversion it would sort before the letters).

  * MY_CS_IDENT_CASEUP_CI - for simple 8bit case insensitive collations.
    It means that this collation has no irregularities on the IDENT
    repertoire (but can have irregularities say on punctuation).

  * MY_CS_ASCII_STD_UCA - for UCA collations.
    It means that a UCA collation does not reorder ASCII characters.

- strings/conf_to_src.c was modified to detect and print
  MY_CS_ASCII_CASEUP_CI and MY_CS_IDENT_CASEUP_CI flags.

- strings/ctype-extra.c was regenerated with new flags.

- Adding a number of MTR tests in
  plugin/func_test/mysql-test/func_test/.
  They display a tailoring by collation name and repertoire as
  returned by:
    cs->coll->tailoring(cs, some_repertoire)
  A dynamically linked plugin function collation_tailoring() was added
  for the purpose of these tests. Tailorings do not depend on PAD/NOPAD
  and on the UCA version, so the function shows "[nopad]" for NOPAD
  collations, and "[version X.Y.Z]" for UCA collations on the ASCII
  repertoire. The version is formatted by a new class UCAVersion,
  which uses a new method CharBuffer::append_uint8().

- Adding the test plugin/func_test/mysql-test/func_test/
  ctype_caseup_ci_verify.test. It checks the order of ASCII characters
  for all collations declared "caseup CI" on the ASCII or IDENT range.

- Adding the test plugin/func_test/mysql-test/func_test/
  collation_tailoring_args.test. It checks the arguments of the plugin
  function collation_tailoring(): the repertoire passed by name or by
  number, and the NULL results for unknown names.

- Adding the test plugin/func_test/mysql-test/func_test/ctype_ldml.test.
  It displays the tailorings of collations loaded at runtime from
  Index.xml: simple 8bit collations (flags detected from the weights)
  and Unicode collations (LDML rules which do or do not reorder ASCII).

- Adding a number of MTR tests mysql-test/main/ctype_xxx_tailoring.test
  They display two-dimensional charts showing which collations
  are compatible on which repertoires.
  Tests for DATE_FORMAT() with an explicit locale were added to
  ctype_big5_tailoring.test, ctype_cp932_tailoring.test and
  ctype_latin5_tailoring.test, which also has the chart for latin5.

- Adding a number of MTR tests in the form of the originally
  reported script for various collations:

    SELECT Insert_priv FROM mysql.user WHERE Insert_priv='...';

  The repeated blocks of queries are shared through
  mysql-test/include/
  ctype_tailoring_0{1,2,3}_mysql_user_Insert_priv*.inc.

- Adding tests to ctype_cp932.test checking that ExtractValue() and
  UpdateXML() still raise "Illegal mix of collations".
bsrikanth-mariadb
MDEV-41243, MDEV-41273 Don't freeze mem_root after engine pushdown

When a statement is pushed down to an engine, the once-per-statement
part of the optimization (e.g. the first_cond_optimization part of
JOIN::optimize(), or the whole of JOIN::optimize() for a pushed down
UNION) is skipped. Re-executing the same prepared statement or stored
routine without pushdown then allocated from a mem_root already marked
ROOT_FLAG_READ_ONLY, hitting an assertion in alloc_root().

Add LEX::dont_freeze_mem_root, set by the pushdown_handler and
derived_handler constructors so it covers every engine implementing
them. Prepared_statement::execute_loop(), sp_head::execute() (through
sp_head::dont_freeze_mem_root) and the SP instruction reparse path no
longer freeze the mem_root of such statements. The flag is only
declared and used in PROTECT_STATEMENT_MEMROOT builds.

Note this disables the mem_root protection for the statement for good,
even if later executions are not pushed down, unless the statement is
re-parsed or re-prepared into a new LEX. For a stored routine one pushed
down instruction disables it for the whole routine.

Add tests for prepared SELECT, UNION, UPDATE and DELETE, prepared
EXPLAIN of a pushed UNION, derived tables, stored procedures,
functions, triggers, cursors and metadata-invalidation reparse, with
pushdown switched on and off. Most prepared statement and routine
tests check the statements received by the remote server, or EXPLAIN,
to verify that pushdown actually happened.
ParadoxV5
MDEV-40271: Disallow `\0` and `\n` in replication connection names

Null bytes conflict with strings null-termination,
\while newline bytes are delimiters in the `multi-master.info` file.
While filename-unsafe chars are canonicalized when mapping names
to filenames; these two still cause trouble in other places.
Therefore, this commit extends `check_master_connection_name()`
to also reject names containing these chars.

This check function is now also used for system variable basenames
(e.g., in `@@replicate_do_db`), which (probably harmlessly)
did not have the existing `MAX_CONNECTION_NAME` check.
Sergei Golubchik
cleanup: change mhnsw_read_first() API to take the value not Item*
Sergei Golubchik
Tree<> - a typesafe wrapper for TREE
Sergei Golubchik
cleanup: KEY::type() and KEY::is_hlindex() methods
Sergei Golubchik
MDEV-25848 Support for Multi-Valued Indexes

* New type of high-level indexes - JSON ARRAY
* sql/index/json.{h,cc}
  * table structure (value varbinary(255), tref varbinary(N), PK(value,tref))
  * insert/update/delete/search works, search uses keyread
  * on insert json array is parsed, values are inserted
    * prefixed by the type tag, only simple types are supported
  * records_in_range works, but only for equalities (MEMBER OF)
* MySQL-compatible syntax `INDEX ((CAST(col AS type ARRAY)))`, but
  * col is a bare name, not an expression, expressions are part of MDEV-35853
  * type is ignored, not stored, also part of MDEV-35853
  * thus `'[1, 0.2, true, "foo", null]'` works fine, type not enforced
  * thus SHOW CREATE TABLE prints `... AS JSON ARRAY`
* `const MEMBER OF (col)` uses index, via
  * range optimizer, QUICK_RANGE_SELECT_ARRAY
  * Item_func_member_of::get_mm_tree()/etc
Jan Lindström
MDEV-40038 tp_foreach() crashes on a dying transaction participant

Bug: tp_foreach() found an engine in hton2plugin[] and passed the
result of plugin_lock() to plugin_hton() without a check. While
reap_plugins() deinitializes an uninstalled engine (PLUGIN_IS_DYING),
the slot stays set until the end of ha_finalize_handlerton(), and
plugin_lock() returns NULL. The server crashed in plugin_hton(), for
example in RESET MASTER through ha_commit_checkpoint_request().

This is a regression from aed5928207a, which changed plugin_foreach()
to tp_foreach() and lost the PLUGIN_IS_READY state mask.

Fix: add plugin_lock_ready(), which locks only a PLUGIN_IS_READY plugin
and reports under LOCK_plugin whether a failed lock was for a READY
plugin (out of memory in debug builds). tp_foreach() skips a plugin that
is not READY and returns an error for a READY plugin that cannot be
locked. An uninstalled but busy engine (PLUGIN_IS_DELETED) is not
visited, as with plugin_foreach() before.

The test uses a DEBUG_SYNC point in ha_finalize_handlerton() to run
RESET MASTER while an engine is deinitialized.

Assisted-by: https://mariadb.org/governance/governance-ai-policy/
drrtuy
MDEV-40957 replication from InnoDB into DuckDB works using FULL mode only.
drrtuy
MDEV-40957 DuckDB engine now returns maximum cost for unimplemented index handler operations effectively disabling them.
drrtuy
MDEV-40957 read-free DML on the slave for DuckDB.
drrtuy
MDEV-40957 mixed-mode replication UPDATE/DELETE for DuckDB.
Yuchen Pei
MDEV-41366 Check prefix key match in iterated partition index scan

The idea of Case 2 in can_skip_merging_scans is that when partition
column key prefix is fixed, and the infix that is also the "partition
by range" column, we can scan each partition in order. This relies on
accurately returning EOF when a partition has no (more) matching rows.
It is possible to have no more matching rows at the first index read
of a partition in a non-exact access, in which case we need to check
the prefix match to rule out false positives and return EOF correctly.

This change by itself would result in incomplete results, if Case 2
incorrectly determines the scan can iterate over partitions in order
in AFTER_KEY and BEFORE_KEY reads. These reads when applied to a
partition column key prefix, necessarily means seeking a different
prefix value. This could result in false negatives. To that end, we
strengthen the checks in Case 2 so that AFTER_KEY and BEFORE_KEY reads
require both the prefix and the partitioned column itself in the
keypart_map to ensure correctness.

For example, this check disqualifies the index_read_map call in an
existing loose index scan testcase, causing the index scan to use both
unordered (in read_range_first) and ordered.

Since we have to enumerate all flags, we include a fix of missing
reverse_order for KEY_OR_PREV. The only known usecase is HANDLER ...
READ which is not working correctly and due to be fixed in MDEV-41364.
For consistency with other equally incorrect HANDLER cases involving
other flags, we do not single out KEY_OR_PREV for special
treatment (e.g. return false in can_skip_merging_scans).
drrtuy
MDEV-40957 various fixes for mixed DML path that suffered from uninit bitmaps and wrong SQL statements that failed in DuckDB.
Sergei Golubchik
skip hlindex update if its columns aren't in the write_set
forkfun
MDEV-31527 Flush buffered early messages on early exit

Buffered_logs::flush() prints and frees the buffered early-option
messages. Call it on every early exit, so warnings and errors are
not lost and the buffer is not leaked.

Make the early-option error exit in mysqld_main() unconditional.
Without the performance schema, a failed early option parse left
"----file-marker----" in argv, and it was reported as an unknown
option.

Index.xml parse errors go through error_handler_hook, not
vprint_msg_to_log(), so --validate-config reported
"Configuration is valid." Install a hook that sets
validate_config_has_warnings.

Test cleanup.
Sergei Golubchik
cleanup: move mhnsw behind the hlindex interface

* new hlindex home: sql/index/
* API classes:
  * hlindexton - singleton, global methods. it creates
  * hlindex_share - a physical index (like TABLE_SHARE), it creates
  * hlindex - one index cursor (like handler or TABLE)
* TABLE stores hlindex, TABLE_SHARE stores hlindex_share, no more void*
* move vector_hnsw.h to its new home sql/index, added sql/index/hlindex.h
* MHNSW_Share is stored in mhnsw_share, Search_context in mhnsw_index
* KEY::options() method