mirror of
https://github.com/cmallwitz/Financials-Extension.git
synced 2026-08-24 10:04:10 -05:00
Workaround for Yahoo HTTPS fingerprinting
This commit is contained in:
@@ -1,14 +1,44 @@
|
||||
# Financials-Extension
|
||||
|
||||
Version 3.3.0 includes improved cookie handling and somewhat improved logic to deal with network issues.
|
||||
|
||||
## Overview
|
||||
|
||||
This is a Python based extension for LibreOffice Calc to make market data available in Calc
|
||||
spreadsheets - currently supporting Yahoo's (FX, crypto, equities, indices, futures, options) and Financial Times'
|
||||
(FX, equities, indices, futures) websites using old-fashioned web scraping.
|
||||
|
||||
Starting with version 3.1.0, we received a contribution to get crypto data directly from Coinbase
|
||||
## Latest version vs Yahoo HTTPS fingerprinting
|
||||
|
||||
Latest version 3.8.0 was created to bypass Yahoo's recently adding crazy HTTPS fingerprinting
|
||||
to their website. If you don't use Yahoo, no further changes are required.
|
||||
|
||||
If you do use Yahoo as a source, here is what I had to do to get Yahoo working again on my Ubuntu system:
|
||||
- install system-wide Python module curl_cffi - on my system as root: `pip3 install curl_cffi --upgrade`
|
||||
- download latest binary of [curl-impersonate](https://github.com/lwthiker/curl-impersonate/releases) e.g.
|
||||
libcurl-impersonate-v0.6.1.x86_64-linux-gnu.tar.gz and untar it somewhere
|
||||
|
||||
Now some of these bits need to be loaded/initialised before running LibreOffice: I used the below (adjust your location
|
||||
to libcurl-impersonate-chrome.so) to run LibreOffice Calc directly from command line:
|
||||
|
||||
```
|
||||
LD_PRELOAD=/tmp/curl-impersonate/libcurl-impersonate-chrome.so CURL_IMPERSONATE=chrome101 /usr/lib/libreoffice/program/soffice.bin --calc
|
||||
```
|
||||
|
||||
With this I can see the below in the output from `=GETREALTIME("SUPPORT")` and the examples.ods file from this repo
|
||||
can load data for Yahoo again.
|
||||
```
|
||||
...
|
||||
requests=curl_cffi_0.10.0
|
||||
LD_PRELOAD=/tmp/curl-impersonate/libcurl-impersonate-chrome.so
|
||||
CURL_IMPERSONATE=chrome101
|
||||
curl_version="libcurl/8.1.1 BoringSSL zlib/1.2.11 brotli/1.0.9 nghttp2/1.56.0"
|
||||
```
|
||||
|
||||
Similar things should be possible on Windows - let me know if [this](https://stackoverflow.com/questions/1178257/ld-preload-equivalent-for-windows-to-preload-shared-libraries)
|
||||
is helpful and share your experience.
|
||||
|
||||
Background: for a normal Python script just installing curl_cffi is enough to bypass Yahoo's HTTPS fingerprinting.
|
||||
Because LibreOffice is loading the stock curl library before executing the extension code directly, the above hack
|
||||
is required. Unless someone tells me otherwise...
|
||||
|
||||
### Feedback requested:
|
||||
|
||||
@@ -31,7 +61,8 @@ Getting data should be as simple as having this in a cell:
|
||||
Codes 21 and 90 stand for "last price" and "close" (see below), respectively.
|
||||
Only Yahoo has historic data available.
|
||||
|
||||
There is a file **examples.ods** there too with usage examples and possible arguments to functions.
|
||||
There is a file **examples.ods** in the Release area too with usage examples
|
||||
and possible arguments to functions.
|
||||
|
||||
You have to check the respective websites to work out what symbol is the right one for you. Make sure today or the date
|
||||
requested is a trading day (exchange is not closed). If a website doesn't have
|
||||
@@ -129,18 +160,19 @@ On my system (Ubuntu) I installed packages: libreoffice-dev libreoffice-java-com
|
||||
|
||||
cd ~/tech/IdeaProjects/Financials-Extension/
|
||||
|
||||
python3 -m unittest discover src
|
||||
\# Assuming curl-cffi is installed, LD_PRELOAD is not required here
|
||||
|
||||
CURL_IMPERSONATE=chrome101 python3 -m unittest discover src
|
||||
|
||||
\# This builds file **Financials-Extension.oxt**
|
||||
|
||||
./compile.sh
|
||||
|
||||
### Tested with:
|
||||
- Windows 10 / LibreOffice Calc 7.1.2.2 / Python 3.8.8
|
||||
- Ubuntu 22.04.1 / LibreOffice Calc 7.3.7.2 / Python 3.10.6
|
||||
- MacOS 10.15.7 / LibreOffice Calc 7.2.0.4 / Python 3.8.10
|
||||
- Ubuntu 22.04.5 / LibreOffice Calc 7.3.7.2 / Python 3.10.12
|
||||
|
||||
(Previous versions)
|
||||
(Previously)
|
||||
- Windows 10 / LibreOffice Calc 7.1.2.2 / Python 3.8.8
|
||||
- MacOS 10.15.7 / LibreOffice Calc 7.2.0.4 / Python 3.8.10
|
||||
- Debian 10.3 / LibreOffice Calc 6.1.5.2 / Python 3.7.3
|
||||
- Ubuntu 20.04.5 / LibreOffice Calc 6.4.7.2 / Python 3.8.10
|
||||
- Ubuntu 18.04.5 / LibreOffice Calc 6 / Python 3.6.9
|
||||
|
||||
+2
-2
@@ -67,7 +67,7 @@ wget "https://files.pythonhosted.org/packages/ec/57/56b9bcc3c9c6a792fcbaf139543c
|
||||
unzip $TMPFILE dateutil/\* -d "${PWD}"/build/
|
||||
rm $TMPFILE
|
||||
|
||||
wget "https://files.pythonhosted.org/packages/9c/3d/a121f284241f08268b21359bd425f7d4825cffc5ac5cd0e1b3d82ffd2b10/pytz-2024.1-py2.py3-none-any.whl" -O $TMPFILE
|
||||
wget "https://files.pythonhosted.org/packages/81/c4/34e93fe5f5429d7570ec1fa436f1986fb1f00c3e0f43a589fe2bbcd22c3f/pytz-2025.2-py2.py3-none-any.whl" -O $TMPFILE
|
||||
unzip $TMPFILE pytz/\* -d "${PWD}"/build/
|
||||
rm $TMPFILE
|
||||
|
||||
@@ -76,7 +76,7 @@ unzip $TMPFILE pyparsing.py -d "${PWD}"/build/
|
||||
rm $TMPFILE
|
||||
|
||||
# Windows LibreOffice 7.1 Python is missing this...
|
||||
wget "https://files.pythonhosted.org/packages/d9/5a/e7c31adbe875f2abbb91bd84cf2dc52d792b5a01506781dbcf25c91daf11/six-1.16.0-py2.py3-none-any.whl" -O $TMPFILE
|
||||
wget "https://files.pythonhosted.org/packages/b7/ce/149a00dd41f10bc29e5921b496af8b574d8413afcd5e30dfa0ed46c2cc5e/six-1.17.0-py2.py3-none-any.whl" -O $TMPFILE
|
||||
unzip $TMPFILE six.py -d "${PWD}"/build/
|
||||
rm $TMPFILE
|
||||
|
||||
|
||||
Binary file not shown.
+58
-139
@@ -8,18 +8,11 @@
|
||||
# version 3 of the License, or (at your option) any later version.
|
||||
|
||||
|
||||
import codecs
|
||||
import gzip
|
||||
import logging
|
||||
import os
|
||||
import pathlib
|
||||
import random
|
||||
import select
|
||||
import urllib.request
|
||||
import urllib.parse
|
||||
from http import cookiejar
|
||||
from http.client import HTTPConnection, HTTPSConnection, HTTPException
|
||||
|
||||
from importlib import util
|
||||
from datacode import Datacode
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
@@ -28,12 +21,24 @@ logger = logging.getLogger(__name__)
|
||||
# logger.setLevel(logging.DEBUG)
|
||||
|
||||
|
||||
class RedirectException(HTTPException):
|
||||
def __init__(self, location):
|
||||
self.location = location
|
||||
curl_cffi_present = not util.find_spec("curl_cffi") is None
|
||||
requests_present = not util.find_spec("requests") is None
|
||||
|
||||
if curl_cffi_present:
|
||||
logger.debug("Importing curl_cffi...")
|
||||
from curl_cffi import requests, __version__ as requests_version, __name__ as requests_name
|
||||
elif requests_present:
|
||||
logger.debug("Importing requests...")
|
||||
import requests
|
||||
requests_version = requests.__version__
|
||||
requests_name = requests.__name__
|
||||
else:
|
||||
raise Exception("Neither curl_cffi nor requests found.")
|
||||
|
||||
# import requests
|
||||
|
||||
|
||||
class HttpException(HTTPException):
|
||||
class HttpException(Exception):
|
||||
def __init__(self, url, response):
|
||||
self.url = url
|
||||
self.response = response
|
||||
@@ -45,150 +50,63 @@ class HttpException(HTTPException):
|
||||
return f"url='{self.url}' status='{self.response}'"
|
||||
if self.response.headers:
|
||||
h = '\n'.join(sorted(self.response.headers.__str__().splitlines(), key=lambda l: l.lower()))
|
||||
return f"url='{self.url}' status={self.response.status} reason='{self.response.reason}'{h}\n"
|
||||
return f"url='{self.url}' status={self.response.status_code} reason='{self.response.reason}' headers={h}\n"
|
||||
else:
|
||||
return f"url='{self.url}' status={self.response.status} reason='{self.response.reason}'"
|
||||
return f"url='{self.url}' status={self.response.status_code} reason='{self.response.reason}'"
|
||||
|
||||
|
||||
class BaseClient:
|
||||
def __init__(self):
|
||||
self.connections = {}
|
||||
self.cookies = cookiejar.CookieJar()
|
||||
self.last_url = None
|
||||
self.redirect_count = 0 # will be set later
|
||||
self.redirect_count = 0
|
||||
|
||||
self.basedir = os.path.join(str(pathlib.Path.home()), '.financials-extension')
|
||||
os.makedirs(self.basedir, exist_ok=True)
|
||||
|
||||
user_agents = [
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:129.0) Gecko/20100101 Firefox/129.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:130.0) Gecko/20100101 Firefox/130.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:131.0) Gecko/20100101 Firefox/131.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:132.0) Gecko/20100101 Firefox/132.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:133.0) Gecko/20100101 Firefox/133.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:134.0) Gecko/20100101 Firefox/134.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:135.0) Gecko/20100101 Firefox/135.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:136.0) Gecko/20100101 Firefox/136.0',
|
||||
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:137.0) Gecko/20100101 Firefox/137.0',
|
||||
]
|
||||
|
||||
self.default_headers = {
|
||||
'User-Agent': random.sample(user_agents, 1)[0],
|
||||
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
|
||||
'Accept-Encoding': 'gzip, deflate',
|
||||
'Accept-Language': 'en-US,en;q=0.5',
|
||||
'Connection': 'keep-alive',
|
||||
'Cache-Control': 'max-age=0'
|
||||
}
|
||||
if curl_cffi_present:
|
||||
self.session = requests.Session()
|
||||
if logger.isEnabledFor(logging.DEBUG) and self.session.curl:
|
||||
self.session.curl.debug()
|
||||
else:
|
||||
self.session = requests.Session()
|
||||
self.session.headers.update({'User-Agent': random.sample(user_agents, 1)[0],
|
||||
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
|
||||
'Accept-Encoding': 'gzip, deflate',
|
||||
'Accept-Language': 'en-US,en;q=0.5',
|
||||
'Connection': 'keep-alive',
|
||||
'Cache-Control': 'max-age=0',
|
||||
})
|
||||
|
||||
self.response = None
|
||||
self.session.max_redirects = 5
|
||||
|
||||
def request(self, method: str, url: str, data=None, headers={}, **kwargs):
|
||||
|
||||
_headers = self.default_headers.copy()
|
||||
if headers:
|
||||
for key, value in headers.items():
|
||||
_headers[key] = value
|
||||
|
||||
if method == 'POST' and 'Content-Type' not in _headers:
|
||||
_headers['Content-Type'] = 'application/x-www-form-urlencoded'
|
||||
|
||||
connection = None
|
||||
|
||||
scheme, _, host, path = url.split('/', 3)
|
||||
|
||||
if (scheme, host) in self.connections:
|
||||
connection = self.connections.get((scheme, host))
|
||||
|
||||
if connection and select.select([connection.sock], [], [], 0)[0]:
|
||||
connection.close()
|
||||
connection = None
|
||||
|
||||
if not connection:
|
||||
logger.debug('Creating connection --------------------------------------------------')
|
||||
connection = HTTPConnection(host, **kwargs) if scheme == 'http:' else HTTPSConnection(host, **kwargs)
|
||||
|
||||
logger.debug('Creating request -----------------------------------------------------')
|
||||
logger.debug("%s %s", method, url)
|
||||
|
||||
self.last_url = url
|
||||
|
||||
# generate and add cookie headers
|
||||
request = urllib.request.Request(url)
|
||||
|
||||
self.cookies.add_cookie_header(request)
|
||||
if request.get_header('Cookie'):
|
||||
_headers['Cookie'] = request.get_header('Cookie')
|
||||
|
||||
for key, value in _headers.items():
|
||||
logger.debug('Header: %s=%s', key, value)
|
||||
|
||||
# request
|
||||
connection.request(method, '/' + path, data, _headers)
|
||||
response = connection.getresponse()
|
||||
|
||||
logger.debug('Processing response --------------------------------------------------')
|
||||
|
||||
logger.debug('response.status=%s', response.status)
|
||||
for key, value in response.getheaders():
|
||||
logger.debug('Header: %s=%s', key, value)
|
||||
|
||||
self.cookies.extract_cookies(response, request)
|
||||
self.connections[(scheme, host)] = connection
|
||||
|
||||
return response
|
||||
|
||||
def urlopen(self, url, redirect=True, data=None, headers={}, cookies=[], **kwargs):
|
||||
|
||||
if cookies:
|
||||
for c in cookies:
|
||||
self.cookies.set_cookie(c)
|
||||
def urlopen(self, url, data=None):
|
||||
|
||||
self.last_url = None
|
||||
|
||||
self.response = self.request('POST' if data else 'GET', url, data, headers, **kwargs)
|
||||
text = self.response.read()
|
||||
resp = self.session.request('POST' if data else 'GET', url, data=data)
|
||||
|
||||
# Allow redirects - used by Yahoo for some cookie based consent
|
||||
self.redirect_count = 5
|
||||
if 400 <= resp.status_code < 500:
|
||||
if resp.headers.get('X-Cache') == 'Error from cloudfront':
|
||||
resp = self.session.request('POST' if data else 'GET', url, data=data)
|
||||
|
||||
# (for Yahoo) AWS CloudFront occasionally returns an incorrect, cached error responses
|
||||
# try mitigating by re-requesting straight away
|
||||
if 400 <= self.response.status < 500:
|
||||
if self.response.getheader('X-Cache') == 'Error from cloudfront':
|
||||
self.response = self.request('POST' if data else 'GET', url, data, headers, **kwargs)
|
||||
text = self.response.read()
|
||||
if resp.status_code >= 400:
|
||||
logger.warning("url='%s' status=%s reason='%s' headers=%s", resp.url,
|
||||
resp.status_code, resp.reason,
|
||||
'\n'.join(sorted(resp.headers.__str__().splitlines(), key=lambda l: l.lower())))
|
||||
raise HttpException(url, resp)
|
||||
|
||||
while 300 <= self.response.status < 400 and self.redirect_count >= 0:
|
||||
self.redirect_count = len(resp.history)
|
||||
self.last_url = resp.url
|
||||
|
||||
self.redirect_count -= 1
|
||||
location = str(self.response.getheader('Location'))
|
||||
location = location.replace(' ', '%20') # FT bug workaround - this should not be necessary
|
||||
|
||||
if location and redirect:
|
||||
|
||||
if location.startswith('/'):
|
||||
scheme, _, host, path = url.split('/', 3)
|
||||
location = '{}//{}{}'.format(scheme, host, location)
|
||||
|
||||
self.response = self.request('GET', location, None, headers, **kwargs)
|
||||
text = self.response.read()
|
||||
|
||||
else:
|
||||
raise RedirectException(location)
|
||||
|
||||
if self.response.status >= 400:
|
||||
logger.warning("last_url='%s' status=%s reason='%s' headers=%s", self.last_url, self.response.status,
|
||||
self.response.reason,
|
||||
'\n'.join(sorted(self.response.headers.__str__().splitlines(), key=lambda l: l.lower())))
|
||||
raise HttpException(url, self.response)
|
||||
|
||||
if self.response.getheader('Content-Encoding') == 'gzip':
|
||||
text = gzip.decompress(text)
|
||||
|
||||
content_type = self.response.headers.get_content_charset()
|
||||
if content_type is None:
|
||||
content_type = 'utf-8'
|
||||
text = codecs.decode(text, encoding=content_type, errors='ignore')
|
||||
|
||||
return text
|
||||
return resp.text
|
||||
|
||||
def get_ticker(self):
|
||||
|
||||
@@ -397,10 +315,11 @@ class BaseClient:
|
||||
|
||||
return None
|
||||
|
||||
def version(self):
|
||||
return requests_name + "_" + requests_version
|
||||
|
||||
def curl(self):
|
||||
return curl_version
|
||||
|
||||
def close(self):
|
||||
for connection in self.connections.values():
|
||||
try:
|
||||
connection.close()
|
||||
except BaseException:
|
||||
pass
|
||||
self.connections = {}
|
||||
self.session.close()
|
||||
|
||||
+14
-1
@@ -240,7 +240,8 @@ class FinancialsImpl(unohelper.Base, Financials):
|
||||
if e.tag.endswith('version'):
|
||||
version = e.attrib['value']
|
||||
|
||||
s = 'ctx={}\nid(self)={}\nversion={}\nfile={}\ncwd={}\nhome={}\nuname={}\npid={}\nsys.executable={}\nsys.version={}\nsys.path={}\nlocale={}\ndefaultlocale={}\ndateutil={}\npytz={}\npyparsing={}\nsix={}'.format(
|
||||
s = ('ctx={}\nid(self)={}\nversion={}\nfile={}\ncwd={}\nhome={}\nuname={}\npid={}\nsys.executable={}\nsys.version={}\nsys.path={}\n' +
|
||||
'locale={}\ndefaultlocale={}\ndateutil={}\npytz={}\npyparsing={}\nsix={}\nrequests={}').format(
|
||||
self.ctx,
|
||||
id(self),
|
||||
version,
|
||||
@@ -258,8 +259,20 @@ class FinancialsImpl(unohelper.Base, Financials):
|
||||
pytz.__version__,
|
||||
pyparsing.__version__,
|
||||
six.__version__,
|
||||
self.ft.version()
|
||||
)
|
||||
|
||||
ld_preload = os.environ.get('LD_PRELOAD')
|
||||
if ld_preload:
|
||||
s += f"\nLD_PRELOAD={ld_preload}"
|
||||
|
||||
curl_impersonate = os.environ.get('CURL_IMPERSONATE')
|
||||
if curl_impersonate:
|
||||
s += f"\nCURL_IMPERSONATE={curl_impersonate}"
|
||||
|
||||
if 'curl_cffi' in self.ft.version():
|
||||
s += f"\ncurl_version=\"{self.ft.session.curl.version().decode()}\""
|
||||
|
||||
if datacode:
|
||||
s = '{}\ntype(datacode)={}\nstr(datacode)={}'.format(
|
||||
s,
|
||||
|
||||
@@ -59,7 +59,7 @@ class Coinbase(BaseClient):
|
||||
url = 'https://api.exchange.coinbase.com/products/{}/stats'.format(ticker)
|
||||
|
||||
try:
|
||||
text = self.urlopen(url, redirect=True, data=None, headers=None)
|
||||
text = self.urlopen(url)
|
||||
except BaseException as e:
|
||||
logger.exception("BaseException ticker=%s datacode=%s last_url=%s redirect_count=%s", ticker, datacode, self.last_url, self.redirect_count)
|
||||
del self.realtime[ticker]
|
||||
|
||||
@@ -76,7 +76,7 @@ class FT(BaseClient):
|
||||
url = f'https://markets.ft.com/data/{asset_class}/tearsheet/summary?s={urllib.parse.quote_plus(ticker)}'
|
||||
|
||||
try:
|
||||
text = self.urlopen(url, redirect=True, data=None, headers=None)
|
||||
text = self.urlopen(url)
|
||||
except BaseException as e:
|
||||
logger.exception("BaseException ticker=%s datacode=%s last_url=%s redirect_count=%s", ticker, datacode, self.last_url, self.redirect_count)
|
||||
del self.realtime[ticker]
|
||||
|
||||
@@ -119,7 +119,7 @@ class Yahoo(BaseClient):
|
||||
def handleCookiesAndConsent(self, url, ticker, datacode, html_file):
|
||||
|
||||
try:
|
||||
text = self.urlopen(url, redirect=True)
|
||||
text = self.urlopen(url)
|
||||
except BaseException as e:
|
||||
logger.exception("BaseException (1) ticker=%s datacode=%s last_url=%s redirect_count=%s %s",
|
||||
ticker, datacode, self.last_url, self.redirect_count, e)
|
||||
@@ -155,7 +155,7 @@ class Yahoo(BaseClient):
|
||||
data[d.attrib['name']] = d.attrib['value']
|
||||
|
||||
try:
|
||||
text = self.urlopen(self.last_url, redirect=True, data=urllib.parse.urlencode(data))
|
||||
text = self.urlopen(self.last_url, data=data)
|
||||
except BaseException as e:
|
||||
logger.exception("BaseException (4) ticker=%s datacode=%s last_url=%s redirect_count=%s %s",
|
||||
ticker, datacode, self.last_url, self.redirect_count, e)
|
||||
@@ -196,7 +196,7 @@ class Yahoo(BaseClient):
|
||||
|
||||
if not self.crumb:
|
||||
|
||||
url = 'https://finance.yahoo.com/quote/{}?p={}'.format(ticker, ticker)
|
||||
url = f'https://finance.yahoo.com/quote/{ticker}'
|
||||
text = self.handleCookiesAndConsent(url, ticker, datacode, f'yahoo-{ticker}.html')
|
||||
|
||||
if text is None:
|
||||
|
||||
@@ -14,7 +14,7 @@ import os
|
||||
cur_dir = os.getcwd()
|
||||
|
||||
addin_id = "com.financials.getinfo"
|
||||
addin_version = "3.7.2"
|
||||
addin_version = "3.8.0"
|
||||
addin_displayname = "Financial Market Extension"
|
||||
addin_publisher_link = "https://github.com/cmallwitz/Financials-Extension"
|
||||
addin_publisher_name = "The Publisher"
|
||||
|
||||
+1
-1
@@ -24,7 +24,7 @@ import testutils
|
||||
financials = financials.createInstance(None)
|
||||
|
||||
|
||||
def urlopen_fail(self, url, redirect=True, data=None, headers={}, cookies=[], **kwargs):
|
||||
def urlopen_fail(self, url, data=None):
|
||||
raise baseclient.HttpException(url, 'ERROR: simulated urlopen() failed')
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user