341

submitted 1 year ago by [email protected] to c/[email protected]

37 comments fedilink hide all child comments

cross-posted from: https://lemmy.sdf.org/post/12950329

you are viewing a single comment's thread
view the rest of the comments

[-] [email protected] 31 points 1 year ago

The section about "regular language" is the reason. That's not being cheeky, that's a technical term. It immediately dives into some complex set theory stuff but that's the place to start understanding.

[-] [email protected] -1 points 1 year ago

English isn't a regular language either. So that means you can't use regex to parse text. But everyone does anyway.

[-] [email protected] 14 points 1 year ago

Huh? Show me the regex to parse the English language.

[-] [email protected] 2 points 1 year ago* (last edited 1 year ago)

Parsing text is the reason regex was created!

Page 1, Chapter 1, "Mastering Regular Expressions", Friedl, O'Reilly 1997.

" Introduction to Regular Expressions

Here's the scenario: you're given the job of checking the pages on a web server for doubled words (such as "this this"), a common problem with documents sub- ject to heavy editing. Your job is to create a solution that will:

Accept any number of files to check, report each line of each file that has doubled words, highlight (using standard ANSI escape sequences) each dou- bled word, and ensure that the source filename appears with each line in the report.

Work across lines, even finding situations where a word at the end of one line is repeated at the beginning of the next.

Find doubled words despite capitalization differences, such as with The the, as well as allow differing amounts of whitespace (spaces, tabs, new- lines, and the like) to lie between the words.

Find doubled words even when separated by HTML tags. HTML tags are for marking up text on World Wide Web pages, for example, to make a word bold: it is very very important

That's certainly a tall order! But, it's a real problem that needs to be solved. At one point while working on the manuscript for this book. I ran such a tool on what I'd written so far and was surprised at the way numerous doubled words had crept in. There are many programming languages one could use to solve the problem, but one with regular expression support can make the job substantially easier.

Regular expressions are the key to powerful, flexible, and efficient text processing. Regular expressions themselves, with a general pattern notation almost like a mini programming language, allow you to describe and parse text.. With additional sup- port provided by the particular tool being used, regular expressions can add, remove, isolate, and generally fold, spindle, and mutilate all kinds of text and data.

Chapter 1: Introduction to Regular Expressions "

[-] [email protected] 16 points 1 year ago

There's a difference between 'processing' the text and 'parsing' it. The processing described in the section you posted it fine, and you can manage a similar level of processing on HTML. The tricky/impossible bit is parsing the languages. For instance you can't write a regex that'll relibly find the subject, object and verb in any english sentence, and you can't write a regex that'll break an HTML document down into a hierarchy of tags as regexs don't support counting depth of recursion, and HTML is irregular anyway, meaning it can't be reliably parsed with a regular parser.

[-] [email protected] -2 points 1 year ago

For instance you can’t write a regex that’ll relibly find the subject, object and verb in any english sentence

Identifying parts of speech isn't a requirement of the word parse. That's the linguistic definition. In computer science identifying tokens is parsing.

https://en.m.wikipedia.org/wiki/Parsing

[-] [email protected] 9 points 1 year ago

That's certainly one level of parsing, and sometimes alk you need, but as the article you posted says, it more usually refers to generating a parse tree. To do that in a natural language isn't happening with a regex.

[-] [email protected] 1 points 1 year ago

Thanks for all the explaining. I always wondered why you can't parse HTML since I first saw the Stack Overflow post, when you can take any HTML code you find and write an expression to work against said set of data.

I never understood the word parse to mean understanding and building a structure based on any input.

[-] [email protected] 9 points 1 year ago

None of these examples are for parsing English sentences. They parse completely different formal languages. That it's text is irrelevant, regex usually operates on text.

You cannot write a regex to give you for example "the subject of an English sentence", just as you can't write a regex to give you "the contents of a complete div tag", because neither of those are regular languages (HTML is context-free, not sure about English, my guess is it would be considered recursively enumerable).

You can't even write a regex to just consume <div> repeated exactly n times followed by </div> repeated exactly n times, because that is already a context-free language instead of a regular language, in fact it is the classic example for a minimal context-free language that Wikipedia also uses.

[-] [email protected] -5 points 1 year ago* (last edited 1 year ago)

None of these examples are for parsing English sentences.

Read it again:

"At one point while working on the manuscript for this book. I ran such a tool on what I’d written so far"

The author explicitly stated that he used regex to parse his own book for errors! The example was using regex to parse html.

You can’t even write a regex to just consume repeated exactly n times followed by repeated exactly n times,

Just because regex can't do everything in all cases doesn't mean it isn't useful to parse some html and English text.

It's like screaming, "You can't build an Operating system with C because it doesn't solve the halting problem!"

this post was submitted on 29 Feb 2024

341 points (96.2% liked)

linuxmemes

26644 readers

2384 users here now

Hint: :q!

Sister communities:

Community rules (click to expand)

1. Follow the site-wide rules

Instance-wide TOS: https://legal.lemmy.world/tos/
Lemmy code of conduct: https://join-lemmy.org/docs/code_of_conduct.html

2. Be civil

Understand the difference between a joke and an insult.

Do not harrass or attack users for any reason. This includes using blanket terms, like "every user of thing".

Don't get baited into back-and-forth insults. We are not animals.

Leave remarks of "peasantry" to the PCMR community. If you dislike an OS/service/application, attack the thing you dislike, not the individuals who use it. Some people may not have a choice.

Bigotry will not be tolerated.

3. Post Linux-related content

Including Unix and BSD.

Non-Linux content is acceptable as long as it makes a reference to Linux. For example, the poorly made mockery of sudo in Windows.

No porn, no politics, no trolling or ragebaiting.

4. No recent reposts

Everybody uses Arch btw, can't quit Vim, <loves/tolerates/hates> systemd, and wants to interject for a moment. You can stop now.

5. 🇬🇧 Language/язык/Sprache

This is primarily an English-speaking community. 🇬🇧🇦🇺🇺🇸

Comments written in other languages are allowed.

The substance of a post should be comprehensible for people who only speak English.

Titles and post bodies written in other languages will be allowed, but only as long as the above rule is observed.

6. (NEW!) Regarding public figures

We all have our opinions, and certain public figures can be divisive. Keep in mind that this is a community for memes and light-hearted fun, not for airing grievances or leveling accusations.

Keep discussions polite and free of disparagement.

We are never in possession of all of the facts. Defamatory comments will not be tolerated.

Discussions that get too heated will be locked and offending comments removed.

Please report posts and comments that break these rules!

Important: never execute code or follow advice that you don't understand or can't verify, especially here. The word of the day is credibility. This is a meme community -- even the most helpful comments might just be shitposts that can damage your system. Be aware, be smart, don't remove France.

founded 2 years ago

MODERATORS

[email protected]