# text

> **Info**
>
> This algorithm is available again, starting with MAGE version 1.22.
> The algorithm was unavailable in MAGE from version 1.14 to version 1.21.

The `text` module offers a toolkit for manipulating strings.

| Trait               | Value      |
| ------------------- | ---------- |
| **Module type**     | util       |
| **Implementation**  | C++        |
| **Parallelism**     | sequential |

## Procedures

### `join()`

Joins all the strings into a single one with the given delimiter between them.

> **Info**
>
> This procedure is equivalent to **apoc.text.join**.

#### Input:

- `subgraph: Graph` (**OPTIONAL**) ➡ A specific subgraph, which is an [object of type Graph](https://memgraph.com/docs/advanced-algorithms/run-algorithms#run-procedures-on-subgraph) returned by the `project()` function, on which the algorithm is run. 
If subgraph is not specified, the algorithm is computed on the entire graph by default.
- `strings: List[string]` ➡ A list of strings to be joined.
- `delimiter: string` ➡ A string to be inserted between the given strings.

#### Output:

- `string: string` ➡ The joined string.

#### Usage:

To join strings, use the following query:

```cypher
CALL text.join(["idora", " ", "ivan", "", "matija"], ",") 
YIELD string 
RETURN string;
```

Result:

```plaintext
+----------------------------+
| string                     |
+----------------------------+
| "idora, ,ivan,,matija"     |
+----------------------------+
```

### `indexOf()`

Finds the index of the first occurrence of a substring within a string, with optional start and end position parameters.

> **Info**
>
> This function is equivalent to **apoc.text.indexOf**.

#### Input:

- `text: string` ➡ The text string to search within.
- `lookup: string` ➡ The substring to search for.
- `from: integer` (default = 0) ➡ The starting position for the search (0-based index).
- `to: integer` (default = -1) ➡ The ending position for the search (-1 means search to the end of the string).

#### Output:

- `integer` ➡ The index of the first occurrence of the substring, or -1 if not found.

#### Usage:

Use the following query to find the position of a substring:

```cypher
RETURN text.indexOf("Hello World!", "World", 0, -1) AS result;
```

Result:

```plaintext
+--------+
| result |
+--------+
| 6      |
+--------+
```

### `regexGroups()`

The procedure returns all matched subexpressions of the regex on the provided
text using the [C++ regex](https://en.cppreference.com/w/cpp/regex) library.

> **Info**
>
> This procedure is equivalent to **apoc.text.regexGroups**.

#### Input:

- `subgraph: Graph` (**OPTIONAL**) ➡ A specific subgraph, which is an [object of type Graph](https://memgraph.com/docs/advanced-algorithms/run-algorithms#run-procedures-on-subgraph) returned by the `project()` function, on which the algorithm is run. 
If subgraph is not specified, the algorithm is computed on the entire graph by default.
- `input: string` ➡ Text that will be searched for regex subexpressions.
- `regex: string` ➡ Regex subexpression searched for in the text. 

#### Output:

- `results: List[List[string]]` ➡ All matched subexpressions. The inner list contains the whole subexpression and tokens matched inside.

#### Usage:

Use the following query to search for expressions: 

```cypher
CALL text.regexGroups("Memgraph: 1\nSQL: 2", "(\\w+): (\\d+)")
YIELD results
RETURN results;
```

Result:

```plaintext
+------------------------------------------------------------+
| results                                                    |
+------------------------------------------------------------+
| [["Memgraph: 1", "Memgraph", "1"], ["SQL: 2", "SQL", "2"]] |
+------------------------------------------------------------+
```

### `format()`

The procedure formats strings using the [C++ fmt library](https://github.com/fmtlib/fmt).

> **Info**
>
> This procedure is equivalent to **apoc.text.format**.

#### Input:

- `subgraph: Graph` (**OPTIONAL**) ➡ A specific subgraph, which is an [object of type Graph](https://memgraph.com/docs/advanced-algorithms/run-algorithms#run-procedures-on-subgraph) returned by the `project()` function, on which the algorithm is run. 
If subgraph is not specified, the algorithm is computed on the entire graph by default.
- `text: string` ➡ Text that needs to be formatted.
- `parameters: string` ➡ Parameters which will be applied to the text.

#### Output:

- `result: string` ➡ Formatted string.

#### Usage:

Use the following queries to insert the parameters to the placeholders in the sentence:

```cypher
CALL text.format("Memgraph is the number {} {} in the world.", [1, "graph database"])
YIELD result
RETURN result;
```

Result:

```plaintext
+---------------------------------------------------------+
| result                                                  |
+---------------------------------------------------------+
| "Memgraph is the number 1 graph database in the world. "|
+---------------------------------------------------------+
```

### `replace()`

Replace each substring of the given string that matches the given regular expression with the given replacement.

> **Info**
>
> This function is equivalent to **apoc.text.replace**.

#### Input:

- `subgraph: Graph` (**OPTIONAL**) ➡ A specific subgraph, which is an [object of type Graph](https://memgraph.com/docs/advanced-algorithms/run-algorithms#run-procedures-on-subgraph) returned by the `project()` function, on which the algorithm is run. 
If subgraph is not specified, the algorithm is computed on the entire graph by default.
- `text: string` ➡ Text that needs to be replaced.
- `regex: string` ➡ Regular expression by which to replace the string.
- `replacement: string` ➡ Target string to replace the matched string.

#### Usage:

Use the following queries to do text replacement:

```cypher
RETURN text.replace('Hello World!', '[^a-zA-Z]', '') AS result;
```

Result:

```plaintext
+--------------+
| result       |
+--------------+
| "HelloWorld" |
+--------------+
```

```cypher
RETURN text.replace('MAGE is a Memgraph Product', 'MAGE', 'GQLAlchemy') AS result;
```

Result:

```plaintext
+------------------------------------+
| result                             |
+---------- -------------------------+
| "GQLAlchemy is a Memgraph Product" |
+------------------------------------+
```

### `regReplace()`

Replace each substring of the given string that matches the given regular expression with the given replacement.

> **Info**
>
> This function is equivalent to **apoc.text.regReplace**.

#### Input:

- `subgraph: Graph` (**OPTIONAL**) ➡ A specific subgraph, which is an [object of type Graph](https://memgraph.com/docs/advanced-algorithms/run-algorithms#run-procedures-on-subgraph) returned by the `project()` function, on which the algorithm is run. 
If subgraph is not specified, the algorithm is computed on the entire graph by default.
- `text: string` ➡ Text that needs to be replaced.
- `regex: string` ➡ Regular expression by which to replace the string.
- `replacement: string` ➡ Target string to replace the matched string.

#### Usage:

Use the following query to do text replacement:

```cypher
RETURN text.regreplace("Memgraph MAGE Memgraph MAGE", "MAGE", "GQLAlchemy") AS output;
```

Result:

```plaintext
+---------------------------------------+
| result                                |
+---------------------------------------+
| "GQLAlchemy MAGE Memgraph GQLAlchemy" |
+---------------------------------------+
```

### `distance()`

Compare the given strings with the Levenshtein distance algorithm.

#### Input:

- `subgraph: Graph` (**OPTIONAL**) ➡ A specific subgraph, which is an [object of type Graph](https://memgraph.com/docs/advanced-algorithms/run-algorithms#run-procedures-on-subgraph) returned by the `project()` function, on which the algorithm is run. 
If subgraph is not specified, the algorithm is computed on the entire graph by default.
- `text1: string` ➡ Source string.
- `text2: string` ➡ Destination string for comparison.

#### Usage:

Use the following query to calculate distance between texts:

```cypher
RETURN text.distance("Levenshtein", "Levenstein") AS result;
```

Result:

```plaintext
+--------+
| result |
+--------+
| 1      |
+--------+
```

### `compare_cleaned()`

Compares two strings for equality after normalizing each one: keeping only ASCII
letters and digits, converting them to lowercase, and dropping everything else
(accents, punctuation, whitespace, and non-ASCII characters).

> **Info**
>
> This function is equivalent to **apoc.text.compareCleaned**.

> **Info**
>
> Normalization is limited to ASCII and performs no Unicode folding, so accented
> and non-ASCII letters are dropped rather than reduced to a base letter. For
> example, `café` cleans to `caf` and is therefore not equal to `cafe`.

#### Input:

- `text1: string` ➡ The first string to normalize and compare. A `null` value results in `false`.
- `text2: string` ➡ The second string to normalize and compare. A `null` value results in `false`.

#### Output:

- `boolean` ➡ `true` if the two normalized strings are equal, and `false` otherwise.

#### Usage:

Use the following query to compare two strings while ignoring case and punctuation:

```cypher
RETURN text.compare_cleaned("Hello, World!", "hello world") AS result;
```

Result:

```plaintext
+--------+
| result |
+--------+
| true   |
+--------+
```
