← journal open source

journal · 2 august 2026

Two merged patches

Notes on two patches merged into projects I had not worked on before. One is a grammar change in Java, the other is a test migration in Python.

Apache ShardingSphere: Doris partition syntax

Project
Apache ShardingSphere, a distributed SQL layer that shards, routes and rewrites queries
Language
Java, ANTLR4 grammar
Change
3 files, 73 lines added, 2 removed
Merged
30 July 2026 into master
Link
apache/shardingsphere #39201

ShardingSphere parses SQL before it can route or rewrite it. The parsing is driven by ANTLR4 grammar files, one set per SQL dialect, which describe every statement the engine accepts. If a dialect has syntax the grammar does not describe, the parser fails on that statement.

Doris is one of the supported dialects, and its grammar had a gap around partitions. An open issue listed statements that were failing to parse. Three kinds of partition syntax were falling through:

-- multi range: generate many partitions from a range and a step
PARTITION BY RANGE (dt) (
  FROM ("2022-01-01") TO ("2022-01-31") INTERVAL 1 DAY
)

-- fixed range: an explicit half open interval
PARTITION p1 VALUES [("2022-01-01"), ("2022-02-01"))

-- per partition properties after the values clause
PARTITION p1 VALUES LESS THAN ("2022-01-01")
  ("storage_policy" = "cooldown", "replication_num" = "1")

None of this was new syntax to the project. Doris already accepted all three forms in ALTER TABLE, and the grammar already had rules for them. They had not been wired into the CREATE TABLE path. So the change mostly reuses rules that already existed.

The main part was splitting the list of partition definitions so an entry could be either an ordinary partition or a multi range expansion:

// before
partitionDefinitions
    : LP_ partitionDefinition (COMMA_ partitionDefinition)* RP_
    ;

// after
partitionDefinitions
    : LP_ partitionDefinitionItem (COMMA_ partitionDefinitionItem)* RP_
    ;

partitionDefinitionItem
    : partitionDefinition | dorisMultiRangePartition
    ;

dorisMultiRangePartition
    : FROM LP_ expr RP_ TO LP_ expr RP_ INTERVAL expr intervalUnit?
    ;

Then partitionDefinition itself gained the fixed range alternative inside its VALUES clause and an optional (LP_ properties RP_) for the per partition property list, reusing the existing properties rule.

repo convention

ShardingSphere marks dialect specific edits inside shared grammar with // DORIS CHANGED BEGIN and // DORIS ADDED BEGIN comment fences. This is not in the contributing guide. I found it by reading the surrounding file. A first revision that used the wrong fence was flagged in review and corrected.

The rest of the diff, 56 of the 73 lines, is test cases: SQL samples plus the expected parse assertions in the project's XML fixture format. A grammar change needs them, because the claim being made is that a statement which used to fail now parses into a specific tree.

Before opening the PR I ran the Doris parser integration suite locally and put the result in the description: 1261 tests, zero failures. The issue also listed a few SQL fragments that were not valid standalone statements, plus one case that already parsed on master. I said so in the description rather than skipping them silently, so it was clear why there were four assertions against a list of seven cases.

Lightly: unittest to pytest

Project
Lightly, a Python library for self supervised learning on images
Language
Python, pytest
Change
7 files, 248 lines added, 301 removed
Merged
28 July 2026 into master
Link
lightly-ai/lightly #2001

This came from an umbrella issue. The maintainers wanted the test suite moved from unittest.TestCase to plain pytest, split into groups so several people could work in parallel without colliding. I took the models group, seven files.

Most of it is mechanical. Drop the TestCase base class, turn setUp into an autouse fixture, replace self.assertEqual with a bare assert, replace assertRaises with pytest.raises, and swap @unittest.skipUnless for @pytest.mark.skipif.

The largest change was to how the suite handled GPU tests. Every device sensitive test existed twice: a real test defaulting to CPU, and a CUDA wrapper that called it back with different arguments.

# before: two methods, one of them a shim
def test_single_projection_head(self, device: str = "cpu", seed: int = 0) -> None:
    ...

@unittest.skipUnless(torch.cuda.is_available(), "skip")
def test_single_projection_head_cuda(self, seed: int = 0) -> None:
    self.test_single_projection_head(device="cuda", seed=seed)

# after: one method, two parametrized runs
@pytest.mark.parametrize("device", ["cpu", "cuda"])
def test_single_projection_head(self, device: str) -> None:
    if device == "cuda" and not torch.cuda.is_available():
        pytest.skip("CUDA not available")
    ...

That is where most of the 301 deleted lines went. The nested subTest loops came out as well, since pytest reports a failing parametrized case on its own.

no behaviour changes

The point of a refactor PR is that nothing changes except the structure. So when a test looked wrong I left it, and when a name was poor I kept it. Any improvement made along the way would be something the reviewer had to check separately, on top of a diff that is otherwise mechanical.

There was also an existing config detail to respect: five of the seven files are excluded from mypy, and two are not. I left the exclusions as they were, matched what the sibling PRs in the same umbrella issue had done, and confirmed in the description that ruff, mypy and pytest were all clean on the directory.

Notes